Back to docs
Analysis & Synthesis

Your Totals Still Add Up: The Four Error Classes That Survive Every Check You Run (2026)

A check can only catch errors that change the thing it looks at. The most common research QA check is satisfied by every possible answer, which makes it no check at all.

Short answer: a check can only catch errors that change the thing it is looking at. The most common quality check in qualitative analysis - confirming that the excerpt counts across your themes add up to the number of excerpts you coded - is satisfied by every possible assignment of those excerpts, including all the wrong ones. It has no diagnostic power at all. Double-entry bookkeeping, which has been running the same kind of check since the fifteenth century, formally names four classes of error that a perfectly balanced ledger cannot see. All four have exact analogues in a coded transcript, and none of them are caught by anything most research teams currently run.

This is a different problem from not checking enough of your data. That question - what fraction should I review - is a question about coverage, and it has its own answer. This is a question about observability: what a check is capable of seeing, even at one hundred percent coverage.

The check that every wrong answer also passes

Start with the arithmetic, because it settles the matter quickly.

Suppose you have 200 coded excerpts distributed across 8 themes. The reconciliation you run is that the theme counts sum to 200. How many possible codings satisfy it?

Every single one. There are 8^200 ways to assign 200 excerpts to 8 themes, and in all of them the counts sum to 200, because each excerpt lands in exactly one theme by construction. The constraint is not a property of correct codings. It is a property of any coding.

A test that passes for every possible input carries zero information about which input you have. It feels like verification because something was computed and the number matched. Nothing was verified.

The contrast with bookkeeping is instructive. A trial balance is genuinely informative because double-entry imposes an external constraint: every transaction is recorded twice, in opposite directions, so the sums must agree for reasons independent of what any individual clerk decided. Take that redundancy away and the check evaporates. Theme counts have no such redundancy. Nobody recorded each excerpt twice in opposite directions, so there is nothing for the total to disagree with.

A balanced ledger is a parity bit

Double-entry is, in engineering terms, a single parity check across the whole ledger. And like any single parity bit, it catches a specific class of corruption and is structurally blind to the rest. Accountants have known precisely which classes for centuries, and the taxonomy is standard in every introductory curriculum.

Omission

A transaction is never recorded at all. The trial balance still agrees, because the missing entry is absent from both sides equally.

Commission

A correct amount is posted to the wrong account within the right category - a payment recorded against telephone expenses instead of utilities. Both sides move by the right amount. The balance holds.

Principle

The transaction is recorded in a fundamentally wrong class - a capital asset purchase treated as an operating expense. The arithmetic is impeccable and the meaning is wrong.

Compensating

Two independent errors of equal and opposite size. Sales of 500 and 1,500 both entered as 1,000. The total is untouched, and the check sees nothing.

The same four classes in a coded transcript

The mapping is exact, which is what makes the borrowed taxonomy worth having.

Bookkeeping classResearch analogueWhy the count check misses it
OmissionAn excerpt nobody coded, or a transcript segment never reviewedIt is absent from the denominator too, so nothing fails to reconcile
CommissionAn excerpt coded to the wrong sub-theme within the right parentParent-level totals are unchanged
PrincipleA feature request coded as a complaintStill exactly one excerpt in exactly one code
CompensatingTwo excerpts swapped between themes A and BBoth theme totals are identical before and after

The compensating case deserves a moment. Take a study where theme A holds 40 excerpts and theme B holds 25. Move six excerpts from A to B, and six from B to A. A still reads 40. B still reads 25. Every count-based check passes at full coverage, your report is unchanged, and yet 15% of theme A and 24% of theme B are now composed of the wrong material. The quotes you pull to illustrate each theme will come from the wrong pile.

What a second pass actually buys

The obvious remedy is to check the work twice. It helps, and the size of the help is measurable - but it is smaller than most people assume, and it has a hard floor.

Garza and colleagues published a systematic review and meta-analysis of data processing error rates in clinical research (Research Square preprint, 2023; PMID 38196643), covering 93 papers published from 1978 to 2008. Their pooled results:

MethodPooled error rate95% CI
Medical record abstraction6.57%5.51, 7.72
Optical scanning0.74%0.21, 1.60
Single-data entry0.29%0.24, 0.35
Double-data entry0.14%0.08, 0.20

Two things in that table matter more than the double-entry line everybody quotes.

First, the reframe. Qualitative coding is not data entry. Typing a number from a form into a field is data entry, at 0.29%. Reading a document, exercising judgement about what it means, and recording a categorical conclusion is medical record abstraction, at 6.57% - which is 22.7 times worse. Thematic coding is an abstraction task in every respect that matters. On 200 excerpts, a 6.57% rate is about 13 miscoded excerpts, and by the taxonomy above most of them will be commission or principle errors that no total can reveal.

Second, the spread. Across those 93 papers, error rates ranged from 2 errors per 10,000 fields to 2,784 errors per 10,000 fields - a factor of nearly 1,400. There is no universal error rate to plan against. There is only your process, measured.

The residual is correlated, not random

Double entry moves 0.29% to 0.14%, removing about 52% of errors. The CAST trial, comparing single and double entry directly, reported an overall rate of "19 per 10,000 fields", with "Error rates were 22 and 15 per 10,000 fields for SE and DE, respectively" - a 32% reduction. And it was not free: "DE took 37% longer than SE, costing each clinic approximately an extra 90 min per month".

So in CAST you paid 37% more time to remove 32% of errors. That is a worse-than-even trade, and the reason is the important part.

A repetition check only works when the two attempts fail independently. Where the source document is smudged, ambiguous, or genuinely open to two readings, both people make the same mistake - and two matching wrong answers look exactly like two matching right ones. The surviving half is not the unlucky remainder. It is a specific, identifiable population: the cases that were hard in a way that pushes both readers the same direction.

This is why doubling up has a floor that adding a third reader barely lowers. The residual is a property of the material, not of the reviewers, and the correct response is to fix the material - clarify the source, tighten the category definitions - rather than to add another pass over it. That is also why Koji treats definition quality, rather than reviewer count, as the lever worth pulling.

Coverage is not observability

It is worth separating two failures that look similar and are not.

Coverage failure: you examined 10% of the data and the defects were in the other 90%. The fix is to examine more. This is the problem that acceptance sampling addresses, and the all-or-none argument covers it properly.

Observability failure: you examined 100% of the data through an instrument that cannot register the defect. Examining it again, or harder, changes nothing.

The theme-count reconciliation is an observability failure. You could check it a thousand times. It will pass every time, on correct and incorrect codings alike, because the quantity it inspects is invariant under exactly the errors you are worried about. More coverage on a blind check is still blindness.

Four checks that can see a swap

Each of these breaks the invariance that makes the count check useless.

  1. Re-derive from source, not from the coding. Sample 20 excerpts, re-read the original transcript span, and ask what code it should get - without seeing the code it has. This is the only check that catches principle errors, because it does not start from the existing classification.
  2. Check the denominator, not just the numerators. Count the transcript segments that received no code. Omission is invisible in theme totals but obvious in the gap between segments available and segments coded.
  3. Reverse the lookup. Instead of asking what is in theme A, pull every excerpt containing a given phrase and ask which themes they landed in. Excerpts that ought to be together and are not will surface immediately, and this is precisely the direction in which a compensating swap becomes visible.
  4. Track composition over time, not totals. If theme A holds 40 excerpts in both the week-one and week-three coding, but only 28 of them are the same excerpts, the total hid a 30% turnover. Comparing identity rather than cardinality is the single highest-yield check on this list, and in Koji it is a comparison of two analysis runs rather than a manual reconciliation.

How Koji handles this

The reason most teams run a blind check is that a real one was prohibitively expensive by hand. Koji changes the cost, not the logic.

  • Every coded item stays attached to its source. Koji analysis produces grounded items, each traceable back to the transcript span it came from, so check 1 above - re-deriving from source - is a click rather than an afternoon of hunting through documents.
  • Stable question IDs carry through the whole pipeline. Identifiers run from the interview plan to the AI interviewer to the analysis to the report aggregation, which means the denominator is always known. You can see which questions a given interview actually spoke to, so omission has somewhere to show up.
  • Re-running analysis is cheap, which makes composition comparison feasible. Comparing which excerpts sit in a theme across two runs is the check that catches compensating swaps, and it is only practical when re-coding costs minutes rather than weeks.
  • Structured questions remove whole error classes rather than detecting them. A single_choice or ranking answer captured as structured data is not abstracted by a human at all, so it carries the 0.29% class of risk rather than the 6.57% class. The structured questions guide covers all six types.
  • Consistent application of one definition. A single analysis pass applies the same theme definition to transcript 1 and transcript 60, which removes the drift that produces commission errors midway through a long coding session.

One honest caveat, because it follows directly from the correlated-residual argument above: an AI analyst is also a single decoder, and running it twice on the same ambiguous excerpt will tend to produce the same answer twice. Automated consistency removes drift-type errors. It does not remove ambiguity in the material, and it should not be reported as though it did.

Common mistakes to avoid

  • Reporting a reconciliation as a quality check. If the numbers could not have failed to add up, saying that they added up tells your stakeholders nothing.
  • Assuming double-coding halves your error rate. CAST measured a 32% reduction for 37% more effort. Budget for the measured number, not the intuitive one.
  • Benchmarking coding accuracy against data-entry rates. Judgement-based abstraction runs an order of magnitude worse - 6.57% against 0.29% - and pretending otherwise understates your error load by more than twenty-fold.
  • Adding a third reviewer to beat the floor. If the residual is correlated, a third pass mostly reproduces the first two. Fix the ambiguity in the definitions or the source instead.
  • Never counting the uncoded. Omission is the easiest of the four classes to detect and the one almost nobody checks for, because it lives outside every total on the report. Koji surfaces it through the question ID trail.

Frequently asked questions

Why does checking that my theme counts add up prove nothing?

Because every possible assignment of excerpts to themes satisfies that constraint. With 200 excerpts and 8 themes there are 8^200 ways to code them, and the counts sum to 200 in all of them, correct and incorrect alike. A test that passes for every possible input cannot distinguish between inputs, so it carries no information about whether your coding is right.

What are compensating errors?

Two independent errors of equal and opposite size that leave the total unchanged. In bookkeeping, sales of 500 and 1,500 both entered as 1,000. In research, two excerpts swapped between two themes. Both theme counts read exactly what they read before, so any check based on totals passes while the composition of both themes is wrong.

How much does double-coding actually reduce errors?

Less than most people expect. The CAST trial measured 22 versus 15 errors per 10,000 fields for single versus double entry, a 32% reduction, while double entry took 37% longer. A pooled meta-analysis of 93 papers found 0.29% for single entry and 0.14% for double, about 52%. The gap between those figures is itself informative: the benefit depends entirely on whether your two passes fail independently.

Why does a second reviewer not catch the remaining errors?

Because a repetition check only works when the two attempts fail independently. Where the source material is genuinely ambiguous, both reviewers tend to make the same misreading, and two matching wrong answers are indistinguishable from two matching right ones. The surviving errors are not a random remainder; they are the specific cases that push both readers the same way.

Is this the same as saying I should check more of my data?

No, and the distinction matters. Checking more data fixes a coverage problem, where the defects were in the portion you did not look at. This is an observability problem: the defect does not change the quantity your check inspects, so examining 100% of the data through that check still finds nothing. More coverage on a blind check is still blindness.

How does Koji make a real check practical?

Every coded item stays linked to the transcript span it came from, so re-deriving a code from source is immediate rather than a manual hunt. Stable question IDs run from the interview plan through analysis to reporting, which makes uncoded material visible, and cheap re-analysis makes it feasible to compare which excerpts sit in a theme across two runs rather than just how many.

Related Resources

Related Articles

Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement

Learn how to measure inter-rater (intercoder) reliability in qualitative research using Cohen's kappa and Krippendorff's alpha, what thresholds count as reliable, and how AI-native tools make consistent coding the default.

Missing Answers vs. Wrong Answers: Why a Blank Is Worth Twice a Confident Guess (2026)

A missing answer and a wrong answer are not two grades of the same problem. Error-correcting codes price them differently, at exactly two to one, and that ratio should change how you design questions.

Poka-Yoke for Research: Mistake-Proofing Studies Before Errors Become Findings (2026)

How to apply poka-yoke (mistake-proofing) from the Toyota Production System to surveys, interviews and analysis, with control vs warning devices and a catalogue of research error-proofing.

How to Build a Qualitative Research Codebook (With Examples and Templates)

A qualitative codebook is the rulebook for how you code your data — code names, definitions, inclusion criteria, examples, and exceptions. Done well, it makes coding consistent across analysts. Done badly, it produces findings nobody can defend.

You Cannot Spot-Check Your Way to Data Quality: The All-or-None Rule for Research QA

A ten-item spot check accepts a 5 percent defective batch 59.9 percent of the time. Deming's all-or-none rule says inspect nothing or inspect everything, and sampling is optimal essentially never.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.