{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-24T11:22:56.121Z"},"content":[{"type":"documentation","id":"2b7d0c20-2b41-4f03-9bf3-9d6fc52aaf4d","slug":"offsetting-errors-aggregate-checks-research","title":"Your Totals Still Add Up: The Four Error Classes That Survive Every Check You Run (2026)","url":"https://www.koji.so/docs/offsetting-errors-aggregate-checks-research","summary":"A check can only catch errors that change the quantity it inspects. Confirming theme counts sum to the excerpt total is satisfied by all 8^200 possible codings, so it has zero diagnostic power. Double-entry bookkeeping names four error classes invisible to a balanced ledger - omission, commission, principle, compensating - and all four map exactly onto coded transcripts. Measured rates: medical record abstraction 6.57 percent versus single-data entry 0.29 percent, and double entry removes only 32 percent of errors for 37 percent more time because the residual is correlated.","content":"**Short answer:** a check can only catch errors that change the thing it is looking at. The most common quality check in qualitative analysis - confirming that the excerpt counts across your themes add up to the number of excerpts you coded - is satisfied by **every possible assignment of those excerpts**, including all the wrong ones. It has no diagnostic power at all. Double-entry bookkeeping, which has been running the same kind of check since the fifteenth century, formally names four classes of error that a perfectly balanced ledger cannot see. All four have exact analogues in a coded transcript, and none of them are caught by anything most research teams currently run.\n\nThis is a different problem from not checking enough of your data. That question - what fraction should I review - is a question about coverage, and it has its own answer. This is a question about **observability**: what a check is capable of seeing, even at one hundred percent coverage.\n\n### The check that every wrong answer also passes\n\nStart with the arithmetic, because it settles the matter quickly.\n\nSuppose you have 200 coded excerpts distributed across 8 themes. The reconciliation you run is that the theme counts sum to 200. How many possible codings satisfy it?\n\nEvery single one. There are `8^200` ways to assign 200 excerpts to 8 themes, and in all of them the counts sum to 200, because each excerpt lands in exactly one theme by construction. The constraint is not a property of *correct* codings. It is a property of *any* coding.\n\nA test that passes for every possible input carries zero information about which input you have. It feels like verification because something was computed and the number matched. Nothing was verified.\n\nThe contrast with bookkeeping is instructive. A trial balance is genuinely informative because double-entry imposes an **external** constraint: every transaction is recorded twice, in opposite directions, so the sums must agree for reasons independent of what any individual clerk decided. Take that redundancy away and the check evaporates. Theme counts have no such redundancy. Nobody recorded each excerpt twice in opposite directions, so there is nothing for the total to disagree with.\n\n### A balanced ledger is a parity bit\n\nDouble-entry is, in engineering terms, a single parity check across the whole ledger. And like any single parity bit, it catches a specific class of corruption and is structurally blind to the rest. Accountants have known precisely which classes for centuries, and the taxonomy is standard in every introductory curriculum.\n\n### Omission\n\nA transaction is never recorded at all. The trial balance still agrees, because the missing entry is absent from both sides equally.\n\n### Commission\n\nA correct amount is posted to the wrong account within the right category - a payment recorded against telephone expenses instead of utilities. Both sides move by the right amount. The balance holds.\n\n### Principle\n\nThe transaction is recorded in a fundamentally wrong class - a capital asset purchase treated as an operating expense. The arithmetic is impeccable and the meaning is wrong.\n\n### Compensating\n\nTwo independent errors of equal and opposite size. Sales of 500 and 1,500 both entered as 1,000. The total is untouched, and the check sees nothing.\n\n### The same four classes in a coded transcript\n\nThe mapping is exact, which is what makes the borrowed taxonomy worth having.\n\n| Bookkeeping class | Research analogue | Why the count check misses it |\n| --- | --- | --- |\n| Omission | An excerpt nobody coded, or a transcript segment never reviewed | It is absent from the denominator too, so nothing fails to reconcile |\n| Commission | An excerpt coded to the wrong sub-theme within the right parent | Parent-level totals are unchanged |\n| Principle | A feature request coded as a complaint | Still exactly one excerpt in exactly one code |\n| Compensating | Two excerpts swapped between themes A and B | Both theme totals are identical before and after |\n\nThe compensating case deserves a moment. Take a study where theme A holds 40 excerpts and theme B holds 25. Move six excerpts from A to B, and six from B to A. A still reads 40. B still reads 25. Every count-based check passes at full coverage, your report is unchanged, and yet 15% of theme A and 24% of theme B are now composed of the wrong material. The quotes you pull to illustrate each theme will come from the wrong pile.\n\n### What a second pass actually buys\n\nThe obvious remedy is to check the work twice. It helps, and the size of the help is measurable - but it is smaller than most people assume, and it has a hard floor.\n\nGarza and colleagues published a systematic review and meta-analysis of data processing error rates in clinical research (Research Square preprint, 2023; PMID 38196643), covering **93 papers published from 1978 to 2008**. Their pooled results:\n\n| Method | Pooled error rate | 95% CI |\n| --- | --- | --- |\n| Medical record abstraction | 6.57% | 5.51, 7.72 |\n| Optical scanning | 0.74% | 0.21, 1.60 |\n| Single-data entry | 0.29% | 0.24, 0.35 |\n| Double-data entry | 0.14% | 0.08, 0.20 |\n\nTwo things in that table matter more than the double-entry line everybody quotes.\n\n**First, the reframe.** Qualitative coding is not data entry. Typing a number from a form into a field is data entry, at 0.29%. Reading a document, exercising judgement about what it means, and recording a categorical conclusion is **medical record abstraction**, at 6.57% - which is **22.7 times worse**. Thematic coding is an abstraction task in every respect that matters. On 200 excerpts, a 6.57% rate is about **13 miscoded excerpts**, and by the taxonomy above most of them will be commission or principle errors that no total can reveal.\n\n**Second, the spread.** Across those 93 papers, error rates ranged from **2 errors per 10,000 fields to 2,784 errors per 10,000 fields** - a factor of nearly 1,400. There is no universal error rate to plan against. There is only your process, measured.\n\n### The residual is correlated, not random\n\nDouble entry moves 0.29% to 0.14%, removing about 52% of errors. The CAST trial, comparing single and double entry directly, reported an overall rate of \"19 per 10,000 fields\", with \"Error rates were 22 and 15 per 10,000 fields for SE and DE, respectively\" - a 32% reduction. And it was not free: \"DE took 37% longer than SE, costing each clinic approximately an extra 90 min per month\".\n\nSo in CAST you paid 37% more time to remove 32% of errors. That is a worse-than-even trade, and the reason is the important part.\n\nA repetition check only works when the two attempts fail **independently**. Where the source document is smudged, ambiguous, or genuinely open to two readings, both people make the same mistake - and two matching wrong answers look exactly like two matching right ones. The surviving half is not the unlucky remainder. It is a **specific, identifiable population**: the cases that were hard in a way that pushes both readers the same direction.\n\nThis is why doubling up has a floor that adding a third reader barely lowers. The residual is a property of the material, not of the reviewers, and the correct response is to fix the material - clarify the source, tighten the category definitions - rather than to add another pass over it. That is also why Koji treats definition quality, rather than reviewer count, as the lever worth pulling.\n\n### Coverage is not observability\n\nIt is worth separating two failures that look similar and are not.\n\n**Coverage failure**: you examined 10% of the data and the defects were in the other 90%. The fix is to examine more. This is the problem that acceptance sampling addresses, and the [all-or-none argument](/docs/research-quality-inspection-sampling) covers it properly.\n\n**Observability failure**: you examined 100% of the data through an instrument that cannot register the defect. Examining it again, or harder, changes nothing.\n\nThe theme-count reconciliation is an observability failure. You could check it a thousand times. It will pass every time, on correct and incorrect codings alike, because the quantity it inspects is invariant under exactly the errors you are worried about. More coverage on a blind check is still blindness.\n\n### Four checks that can see a swap\n\nEach of these breaks the invariance that makes the count check useless.\n\n1. **Re-derive from source, not from the coding.** Sample 20 excerpts, re-read the original transcript span, and ask what code it should get - without seeing the code it has. This is the only check that catches principle errors, because it does not start from the existing classification.\n2. **Check the denominator, not just the numerators.** Count the transcript segments that received **no** code. Omission is invisible in theme totals but obvious in the gap between segments available and segments coded.\n3. **Reverse the lookup.** Instead of asking *what is in theme A*, pull every excerpt containing a given phrase and ask which themes they landed in. Excerpts that ought to be together and are not will surface immediately, and this is precisely the direction in which a compensating swap becomes visible.\n4. **Track composition over time, not totals.** If theme A holds 40 excerpts in both the week-one and week-three coding, but only 28 of them are the same excerpts, the total hid a 30% turnover. Comparing identity rather than cardinality is the single highest-yield check on this list, and in Koji it is a comparison of two analysis runs rather than a manual reconciliation.\n\n## How Koji handles this\n\nThe reason most teams run a blind check is that a real one was prohibitively expensive by hand. Koji changes the cost, not the logic.\n\n- **Every coded item stays attached to its source.** Koji analysis produces grounded items, each traceable back to the transcript span it came from, so check 1 above - re-deriving from source - is a click rather than an afternoon of hunting through documents.\n- **Stable question IDs carry through the whole pipeline.** Identifiers run from the interview plan to the AI interviewer to the analysis to the report aggregation, which means the denominator is always known. You can see which questions a given interview actually spoke to, so omission has somewhere to show up.\n- **Re-running analysis is cheap, which makes composition comparison feasible.** Comparing which excerpts sit in a theme across two runs is the check that catches compensating swaps, and it is only practical when re-coding costs minutes rather than weeks.\n- **Structured questions remove whole error classes rather than detecting them.** A `single_choice` or `ranking` answer captured as structured data is not abstracted by a human at all, so it carries the 0.29% class of risk rather than the 6.57% class. The [structured questions guide](/docs/structured-questions-guide) covers all six types.\n- **Consistent application of one definition.** A single analysis pass applies the same theme definition to transcript 1 and transcript 60, which removes the drift that produces commission errors midway through a long coding session.\n\nOne honest caveat, because it follows directly from the correlated-residual argument above: an AI analyst is also a single decoder, and running it twice on the same ambiguous excerpt will tend to produce the same answer twice. Automated consistency removes drift-type errors. It does not remove ambiguity in the material, and it should not be reported as though it did.\n\n## Common mistakes to avoid\n\n- **Reporting a reconciliation as a quality check.** If the numbers could not have failed to add up, saying that they added up tells your stakeholders nothing.\n- **Assuming double-coding halves your error rate.** CAST measured a 32% reduction for 37% more effort. Budget for the measured number, not the intuitive one.\n- **Benchmarking coding accuracy against data-entry rates.** Judgement-based abstraction runs an order of magnitude worse - 6.57% against 0.29% - and pretending otherwise understates your error load by more than twenty-fold.\n- **Adding a third reviewer to beat the floor.** If the residual is correlated, a third pass mostly reproduces the first two. Fix the ambiguity in the definitions or the source instead.\n- **Never counting the uncoded.** Omission is the easiest of the four classes to detect and the one almost nobody checks for, because it lives outside every total on the report. Koji surfaces it through the question ID trail.\n\n## Frequently asked questions\n\n### Why does checking that my theme counts add up prove nothing?\n\nBecause every possible assignment of excerpts to themes satisfies that constraint. With 200 excerpts and 8 themes there are 8^200 ways to code them, and the counts sum to 200 in all of them, correct and incorrect alike. A test that passes for every possible input cannot distinguish between inputs, so it carries no information about whether your coding is right.\n\n### What are compensating errors?\n\nTwo independent errors of equal and opposite size that leave the total unchanged. In bookkeeping, sales of 500 and 1,500 both entered as 1,000. In research, two excerpts swapped between two themes. Both theme counts read exactly what they read before, so any check based on totals passes while the composition of both themes is wrong.\n\n### How much does double-coding actually reduce errors?\n\nLess than most people expect. The CAST trial measured 22 versus 15 errors per 10,000 fields for single versus double entry, a 32% reduction, while double entry took 37% longer. A pooled meta-analysis of 93 papers found 0.29% for single entry and 0.14% for double, about 52%. The gap between those figures is itself informative: the benefit depends entirely on whether your two passes fail independently.\n\n### Why does a second reviewer not catch the remaining errors?\n\nBecause a repetition check only works when the two attempts fail independently. Where the source material is genuinely ambiguous, both reviewers tend to make the same misreading, and two matching wrong answers are indistinguishable from two matching right ones. The surviving errors are not a random remainder; they are the specific cases that push both readers the same way.\n\n### Is this the same as saying I should check more of my data?\n\nNo, and the distinction matters. Checking more data fixes a coverage problem, where the defects were in the portion you did not look at. This is an observability problem: the defect does not change the quantity your check inspects, so examining 100% of the data through that check still finds nothing. More coverage on a blind check is still blindness.\n\n### How does Koji make a real check practical?\n\nEvery coded item stays linked to the transcript span it came from, so re-deriving a code from source is immediate rather than a manual hunt. Stable question IDs run from the interview plan through analysis to reporting, which makes uncoded material visible, and cheap re-analysis makes it feasible to compare which excerpts sit in a theme across two runs rather than just how many.\n\n## Related Resources\n\n- [You Cannot Spot-Check Your Way to Data Quality](/docs/research-quality-inspection-sampling) - the coverage half of this problem, and why partial inspection fails on its own terms.\n- [Missing Answers vs. Wrong Answers](/docs/missing-answers-vs-wrong-answers-research) - why an error you can locate costs half as much as one you cannot.\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) - how coding agreement is measured once you have two passes to compare.\n- [How to Build a Qualitative Research Codebook](/docs/qualitative-research-codebook) - the definitions whose ambiguity produces the correlated residual described here.\n- [Poka-Yoke for Research](/docs/poka-yoke-mistake-proofing-research) - designing studies so the error cannot be made, rather than detecting it afterwards.\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and why structured answers skip the abstraction step entirely.","category":"Analysis & Synthesis","lastModified":"2026-09-24T03:24:53.203191+00:00","metaTitle":"Offsetting Errors: Why Your Research Checks Pass Anyway (2026)","metaDescription":"The theme-count check passes for every possible coding. Four error classes from double-entry bookkeeping that your research QA cannot see.","keywords":["offsetting errors","compensating errors","research quality assurance","coding error detection","double data entry error rate","thematic analysis accuracy","data verification","qualitative coding errors"],"aiSummary":"A check can only catch errors that change the quantity it inspects. Confirming theme counts sum to the excerpt total is satisfied by all 8^200 possible codings, so it has zero diagnostic power. Double-entry bookkeeping names four error classes invisible to a balanced ledger - omission, commission, principle, compensating - and all four map exactly onto coded transcripts. Measured rates: medical record abstraction 6.57 percent versus single-data entry 0.29 percent, and double entry removes only 32 percent of errors for 37 percent more time because the residual is correlated.","aiPrerequisites":["Familiarity with qualitative coding or thematic analysis","Basic understanding of research quality assurance"],"aiLearningOutcomes":["Tell an observability failure apart from a coverage failure","Identify the four error classes invisible to aggregate checks","Benchmark coding accuracy against abstraction rather than data-entry error rates","Design checks that can detect a compensating swap"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}