{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-20T13:47:28.739Z"},"content":[{"type":"documentation","id":"35e5e5c7-abf0-41ca-903a-ee562191471f","slug":"research-quality-inspection-sampling","title":"You Cannot Spot-Check Your Way to Data Quality: The All-or-None Rule for Research QA","url":"https://www.koji.so/docs/research-quality-inspection-sampling","summary":"Shows why sampling review of research outputs is a dominated policy. Gives computed operating characteristic curves for common spot-check plans, states Deming's all-or-none inspection rule and how to compute the break-even defect rate, explains why reviewer accuracy falls as the base rate of defects rises, and gives an alternative that moves spend upstream.","content":"**Bottom line up front:** The reflexive fix for a research process that keeps letting bad data through is a spot-check: review a handful of transcripts, sample a few coded excerpts, sign off. Quality engineering has known since the 1950s that the sample review is the one option that is almost never the right one. A spot-check of ten items detects a 5 percent defect rate only 40.1 percent of the time. Deming's all-or-none rule says the cost-minimising policy is a corner solution: inspect nothing, or inspect everything, depending on whether your defect rate is above or below the ratio of the cost of one check to the cost of one escape. Sampling sits at the optimum essentially never.\n\n## The move everyone makes, and the number that kills it\n\nA team measures its [research pipeline yield](/docs/research-pipeline-yield-rolled-throughput), finds it uncomfortable, and does the obvious thing: adds a review gate. Someone will check a sample of transcripts. Someone will spot-check the coding. The reasoning feels unimpeachable. Checking some is better than checking none, and checking all is too expensive.\n\nBoth halves of that sentence are wrong in a way that is exactly computable.\n\nHere is what a single-sample plan actually detects. Each cell is the probability that the plan **accepts** the batch, given the true defect rate. A plan of n=10, c=0 means \"review ten items and reject the batch if you find one or more defects\" - the strictest spot-check most teams would ever run.\n\n| Plan | 1% defective | 2% | 5% | 10% | 20% | 30% |\n| --- | --- | --- | --- | --- | --- | --- |\n| Review 10, reject on 1 | 90.4% | 81.7% | 59.9% | 34.9% | 10.7% | 2.8% |\n| Review 20, reject on 1 | 81.8% | 66.8% | 35.8% | 12.2% | 1.2% | 0.1% |\n| Review 20, reject on 2 | 98.3% | 94.0% | 73.6% | 39.2% | 6.9% | 0.8% |\n| Review 50, reject on 1 | 60.5% | 36.4% | 7.7% | 0.5% | 0.0% | 0.0% |\n\n*Computed from the binomial distribution; each cell is the probability of observing at most c defects in n draws at the stated defect rate.*\n\nRead the top row. A batch of interviews that is 5 percent junk sails through your strict ten-item review 59.9 percent of the time. A batch that is 10 percent junk sails through 34.9 percent of the time. The review is not catching a 10 percent defect rate; it is catching a coin flip's worth of a 10 percent defect rate, and every time it passes, someone writes \"reviewed\" in a status column and the batch is treated as clean.\n\nTurn the same arithmetic the other way: with a 5 percent defect rate, a ten-item spot-check finds at least one defect only 40.1 percent of the time. At 10 percent it finds one 65.1 percent of the time. You would need to be at a 20 percent defect rate before a ten-item check is more likely than not to be genuinely informative, and at 20 percent you did not need a check to know you had a problem.\n\n## The vocabulary that makes this precise\n\nThe field that owns this problem is acceptance sampling, and its terms are worth borrowing because they name things research teams argue about without labels. The NIST *Engineering Statistics Handbook* defines them cleanly:\n\n- **Operating characteristic curve.** \"This curve plots the probability of accepting the lot (Y-axis) versus the lot fraction or percent defectives (X-axis). The OC curve is the primary tool for displaying and investigating the properties of a LASP.\" Every review policy has an OC curve whether or not anyone has drawn it. The table above is one.\n- **Acceptable quality level.** \"The AQL is a percent defective that is the base line requirement for the quality of the producer's product.\" Your research team has an implicit AQL. It has probably never been said out loud.\n- **Lot tolerance percent defective.** \"The LTPD is a designated high defect level that would be unacceptable to the consumer.\"\n- **Producer's risk.** \"the probability, for a given (n,c) sampling plan, of rejecting a lot that has a defect level equal to the AQL.\" This is the cost nobody counts: the good batch you throw away and re-field.\n- **Consumer's risk.** \"the probability, for a given (n,c) sampling plan, of accepting a lot with a defect level equal to the LTPD.\"\n\nThe two risks are the point. A review gate has a false-reject rate as well as a false-accept rate, and the false rejects are expensive in research because re-fielding a study costs weeks. A review step is not a free filter bolted onto the side of the process. It is another stage in the series, with its own yield, and it can lower the total.\n\n## Deming's all-or-none rule\n\nW. Edwards Deming derived the condition under which inspecting at all is worth it, in Chapter 15 of *Out of the Crisis* (MIT CAES, 1986). It is usually called the kp rule or the all-or-none rule, and it is startlingly simple.\n\nLet **k1** be the cost of inspecting one item and **k2** be the cost of letting one defective item through to be discovered downstream. Let **p** be the incoming fraction defective. Then:\n\n- If **p is less than k1/k2**, minimum average cost occurs with **no inspection**.\n- If **p is greater than k1/k2**, minimum average cost occurs with **100 percent inspection**.\n\nThere is no middle. Sampling inspection is optimal only at the exact break-even point, which is a measure-zero coincidence. Every intermediate policy - review a fifth, review a sample, review the ones that look odd - is dominated by one of the two corners.\n\nFor a research team, the ratio is easy to estimate:\n\n| If one check costs | And one escape costs | Break-even defect rate |\n| --- | --- | --- |\n| 1 unit | 20 units | 5.0% |\n| 1 unit | 50 units | 2.0% |\n| 1 unit | 100 units | 1.0% |\n\n*Break-even is k1/k2 by construction.*\n\nNow put real research numbers in. Reading one transcript carefully costs, say, fifteen minutes. A bad interview that reaches synthesis and lands in a readout costs a wrong feature decision, or at minimum an analyst's day plus a credibility hit. If you put k2 at fifty times k1 - conservative for anything feeding a roadmap decision - your break-even defect rate is 2 percent. Almost no research pipeline runs below a 2 percent defect rate at the elicitation stage. Which means the rule's answer for most teams is not \"sample more carefully.\" It is **inspect everything**, and if you cannot afford to inspect everything, that is a statement about your process cost, not a licence to sample.\n\nDeming's own position on what inspection buys you was blunter still. From *Out of the Crisis*, page 29: \"Inspection does not improve the quality, nor guarantee quality. Inspection is too late.\" On the same page he quotes Harold F. Dodge, the Bell Labs statistician who invented acceptance sampling in the first place: \"You can not inspect quality into a product.\" And on page 227: \"Quality can not be inspected into a product or service; it must be built into it.\" The third of Deming's fourteen points is the instruction that follows: \"Cease dependence on inspection to achieve quality. Eliminate the need for inspection on a mass basis by building quality into the product in the first place.\"\n\n## The reviewer has an error rate too, and it gets worse as things get worse\n\nEven 100 percent review is not 100 percent detection, and the way detection degrades is counterintuitive enough that it deserves its own number.\n\nGordon and colleagues, in *Health Expectations* 27(3), 2024, make the point with a worked Bayesian example while analysing fraudulent survey respondents. Suppose a fraud-detection strategy has 90 percent sensitivity and 90 percent specificity, which is far better than most manual review. Their result, stated in the paper:\n\n- At **50 percent** fraud prevalence, \"approximately 90% of responses we determine to be authentic will truly be authentic.\"\n- At **80 percent** prevalence, \"only approximately 69% of responses classified as authentic would be truly authentic.\"\n- At **90 percent** prevalence, \"the proportion of truly authentic responses decreases to 50%.\"\n\nThe detector did not get worse. The base rate did. The same review procedure that is trustworthy on a clean batch becomes a coin flip on a dirty one, which is precisely the situation in which teams lean hardest on it. This is the deep reason inspection cannot substitute for process control: inspection quality is a function of the incoming quality it is supposed to be protecting you from.\n\nThe same arithmetic explains why the two Gordon numbers from the previous article matter so much. Their suspected-fraud rate was 17.4 percent from the initial channel and 83.1 percent after the study was posted to social media. A review policy calibrated on the first channel is nearly useless on the second, and nothing in the review policy itself signals that it has stopped working.\n\n## The sign inversion\n\nThe previous article in this series said: every stage you add lowers the yield of the chain, because stages multiply. The natural inference is that you should add a checking stage to catch what the other stages drop.\n\nThat inference is backwards, and the reason is the whole point of this article. A checking stage is not an exception to the multiplication rule; it is subject to it. It adds its own false-reject rate to the chain, its own delay, and its own escape rate. Adding inspection to a low-yield process gives you a low-yield process that also takes longer and occasionally throws away good work.\n\nThe correct move has the opposite shape. In the yield article, the lever was **measure each stage and fund the minimum**. Here, the lever is **stop measuring harder and change the stage that is producing defects**, because the measurement is itself a sample with an OC curve and its power is worse than your intuition. More checking is not a weaker version of better process. It is a different, dominated policy.\n\n## What to do instead\n\n1. **Compute your break-even.** Estimate k1 (cost of one check) and k2 (cost of one escape) for each stage. The ratio is your threshold defect rate. It takes ten minutes and it usually surprises people.\n2. **Estimate p per stage, once, properly.** Not a spot-check. Take one batch and inspect all of it. You are not doing quality control; you are measuring the process so you never have to guess again.\n3. **If p is above the threshold, do not sample. Fix or automate.** The two ways out of expensive 100 percent inspection are to reduce the defect rate at source, or to make inspection so cheap that k1 collapses and the break-even moves out of reach.\n4. **Mistake-proof at capture, not at review.** A screener question that a fraudulent respondent cannot answer beats any amount of downstream transcript reading, because it moves the defect out of the batch instead of finding it in the batch. See [survey fraud and respondent quality](/docs/survey-fraud-respondent-quality) for the specific mechanisms.\n5. **Keep review for what it is genuinely good at: learning, not filtering.** Deming's objection is to inspection *as a quality strategy*, not to reading your own data. A structured peer review of a study design catches whole classes of error that no sampling plan addresses, and it happens before the defects are created. [The research peer review QA gate](/docs/research-peer-review-qa-gate) is the right shape for this: a gate on the design, not a filter on the output.\n6. **Say the OC curve out loud when someone proposes a sample review.** \"We will check ten\" is a policy with a known false-accept rate. Write it in the doc.\n\n## When 100 percent inspection is right\n\nThe rule cuts both ways, and the \"inspect nothing\" corner is real. If your incoming defect rate is genuinely below k1/k2 - a small internal panel of known customers, a study run with an established screener on a channel you have measured - then reviewing is a net cost. The honest version of that policy is to say so, rather than performing a review that has a 90.4 percent chance of accepting a 1-percent-defective batch and calling it assurance.\n\nBetween the corners, the thing that changes the answer is k1. If checking one item costs almost nothing, the break-even defect rate falls toward zero and 100 percent inspection wins for every realistic p. That is the lever worth pulling, and it is a tooling question rather than a process question.\n\n## How Koji flips the corner solution\n\nKoji's design attacks k1 rather than the sampling plan, which is the only move the rule licenses.\n\n- **Every interview is scored, not a sample.** Koji generates a quality score from 1 to 5 for each completed interview against the study's research goals as it finishes. The inspection rate is 100 percent by construction, so consumer's risk from sampling is zero, and no batch is ever \"reviewed\" in a way that means \"ten of them were.\"\n- **The check is not a separate stage.** Because the score is produced from the same analysis pass that produces the themes, review does not add a handoff to the chain. It costs no additional elapsed time, which is what pushes k1 low enough for the all-or-none rule to land on the \"inspect everything\" corner.\n- **Defects get prevented at capture.** Koji's six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - constrain the answer space where constraint is appropriate, so a whole class of unusable response never enters the batch. See the [structured questions guide](/docs/structured-questions-guide) for how the types map to analysis and reporting.\n- **The AI moderator re-probes instead of accepting a thin answer.** A traditional interview's defects are created live and can only be found later. An AI moderator that follows up on a non-answer removes the defect at the moment it would otherwise be baked in, which is Deming's point about building quality in rather than inspecting it in, implemented literally.\n- **You can still read everything.** Full transcripts remain available, so the 100 percent inspection corner is genuinely available to a human when the decision warrants it, rather than being priced out.\n\nLegacy survey tooling gives you the opposite economics. Fielding is cheap, review is expensive and manual, so every team lands on sampling - the one policy the mathematics rules out.\n\n## Honest objections\n\n**\"Deming's rule assumes a stable process, and ours is not.\"** Correct, and Deming said so: the kp rule applies to a process in statistical control. For an unstable process, the honest reading is worse for sampling, not better, because an unstable p means your sampling plan's OC curve is calibrated to a defect rate that is no longer current. This is exactly the Gordon 17.4 to 83.1 percent case.\n\n**\"We cannot inspect everything, so we have to sample.\"** That is a real constraint, but it should be recorded as a known unmanaged risk rather than as assurance. If you sample ten and accept, write down the probability that you have just accepted a 10 percent defective batch. It is 34.9 percent.\n\n**\"Acceptance sampling is standard practice across whole industries.\"** It was, and the profession has been arguing about it since Mood's theorem and Deming's critique. Acceptance sampling answers \"should I accept this lot at a stated risk,\" which is a supplier-relations question. It does not answer \"how good is my process,\" which is the question research teams are actually asking when they spot-check.\n\n**\"Our reviewers are better than 90 percent sensitivity.\"** Possibly, on clean batches. Ask what their sensitivity is on a batch that is 80 percent bad, then re-read the Gordon numbers. Reviewer performance is not a constant.\n\n## Frequently asked questions\n\n### How many transcripts should we spot-check?\n\nThe honest answer from the mathematics is: none, or all of them. Compute k1/k2 - the cost of one check divided by the cost of one defect escaping - and compare it to your actual defect rate. If your defect rate is higher, review everything. If it is lower, reviewing is a net cost. A partial review is dominated by one of those two policies at essentially every defect rate.\n\n### What does a spot-check of ten actually detect?\n\nNot much at the rates that matter. Reviewing ten items and rejecting on any defect accepts a 5-percent-defective batch 59.9 percent of the time and a 10-percent-defective batch 34.9 percent of the time. Put the other way, at a 5 percent defect rate a ten-item check finds at least one defect only 40.1 percent of the time.\n\n### What is Deming's all-or-none rule?\n\nIt is the cost-minimising inspection policy derived in Chapter 15 of *Out of the Crisis*. With k1 the cost of inspecting one item and k2 the cost of a defect escaping, inspect nothing if the incoming fraction defective is below k1/k2 and inspect everything if it is above. Sampling is optimal only exactly at the break-even point, so in practice it is never the right policy.\n\n### Does that mean peer review of research is a waste of time?\n\nNo, and this is the important distinction. Deming's objection is to inspection as a substitute for process quality. Reviewing a study *design* before fielding prevents defects rather than filtering them, which is the thing he was arguing for. Reviewing a sample of outputs after fielding is the thing he was arguing against.\n\n### Why does a reviewer's accuracy fall when data quality falls?\n\nBecause predictive value depends on base rate, not just on sensitivity and specificity. Gordon and colleagues show that a detector with 90 percent sensitivity and specificity leaves about 90 percent of \"authentic\" classifications truly authentic at 50 percent fraud prevalence, about 69 percent at 80 percent prevalence, and 50 percent at 90 percent prevalence. Your review gets least trustworthy exactly when you need it most.\n\n### What is the producer's risk in a research review gate?\n\nIt is the chance of rejecting and re-fielding a batch that was actually fine. Research teams almost never count this cost, but re-fielding a study is weeks of elapsed time and a delayed decision. A review policy has to be judged on both risks, and tightening the acceptance number to reduce escapes always raises false rejects.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - constraining the answer space so defects are never created\n- [Research Pipeline Yield](/docs/research-pipeline-yield-rolled-throughput) - the series-model arithmetic that makes an extra review stage costly\n- [When No Study Was Wrong](/docs/systemic-research-failure-no-defective-study) - the failures that no amount of inspection can catch\n- [The Research Peer Review QA Gate](/docs/research-peer-review-qa-gate) - review applied to the design, where it prevents rather than filters\n- [Survey Fraud and Respondent Quality](/docs/survey-fraud-respondent-quality) - moving defects out of the batch at the screener\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) - measuring the coding stage rather than sampling it\n- [Research Calibration and Brier Scores](/docs/research-calibration-brier-score) - scoring judgement quality once you have stopped sampling it\n","category":"Research Methods","lastModified":"2026-08-20T03:30:04.019935+00:00","metaTitle":"Spot-Checks Do Not Work: The All-or-None Rule for Research QA","metaDescription":"A ten-item spot check accepts a 5 percent defective batch 59.9 percent of the time. Deming's all-or-none rule and what to do instead of sampling your research data.","keywords":["research data quality","acceptance sampling","operating characteristic curve","Deming all or none rule","research QA","spot check","survey data quality"],"aiSummary":"Shows why sampling review of research outputs is a dominated policy. Gives computed operating characteristic curves for common spot-check plans, states Deming's all-or-none inspection rule and how to compute the break-even defect rate, explains why reviewer accuracy falls as the base rate of defects rises, and gives an alternative that moves spend upstream.","aiPrerequisites":["Understanding that research stages sit in series and their yields multiply"],"aiLearningOutcomes":["Compute the false-accept rate of any spot-check policy","Calculate your own inspect-nothing versus inspect-everything break-even","Explain why a reviewer gets less reliable as data quality falls","Redirect quality spend from review gates to capture-stage prevention"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}