{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-28T00:58:07.399Z"},"content":[{"type":"documentation","id":"dfbe0ef2-3727-4904-8f0b-57620827535a","slug":"zero-numerator-rule-of-three-research","title":"Nobody Mentioned It: What Zero Occurrences Actually Rules Out (2026)","url":"https://www.koji.so/docs/zero-numerator-rule-of-three-research","summary":"Observing zero occurrences in n interviews does not establish a rate of zero; it establishes a 95 percent upper bound of approximately 3/n. Twenty interviews with no mentions are consistent with a 15 percent prevalence. The rule originates with Hanley and Lippman-Hand (JAMA 1983) and is operationalised by Eypasch et al (BMJ 1995) as maximum risk equals 3/n for n above 30. This makes sample size a function of the rate you must rule out rather than of saturation, and it requires a real denominator, which only questions asked of every participant can supply.","content":"**Short answer:** if nobody raised an issue in 20 interviews, you have not learned that the issue is absent. You have learned that its rate is probably below about 15 percent. The correct output of an absence is not zero - it is an upper bound, and the bound is roughly 3 divided by your sample size. Clinical medicine formalised this in 1983 and calls it the rule of three. Most product decisions that begin \"it never came up, so let us deprioritise it\" are quietly assuming a bound they never computed.\n\nThis single move - reporting 3/n instead of zero - is the highest-leverage change most research reports could make, because absences are cited as evidence constantly and almost never quantified.\n\n## The rule of three\n\nWhen you observe zero events in n independent observations, \"the interval from 0 to 3/n is a 95% confidence interval for the rate of occurrences in the population\". Eypasch and colleagues, writing in the BMJ in 1995, state the operational form precisely: \"upper limit of 95% confidence interval = maximum risk = 3/n (for n > 30)\". Their closing advice generalises beyond surgery: \"Doctors and surgeons should keep this simple rule in mind when complication rates of zero are reported in the literature and when they have not (yet) experienced a disastrous complication in a procedure.\" That second clause is the one product teams should read twice: it covers the case where you have personally never seen the failure.\n\nThe canonical worked example: \"a pain-relief drug is tested on 1500 human subjects, and no adverse event is recorded. From the rule of three, it can be concluded with 95% confidence that fewer than 1 person in 500 (or 3/1500) will experience an adverse event.\"\n\nThe rule is not a heuristic somebody invented. It \"can then be derived either from the Poisson approximation to the binomial distribution, or from the formula (1 - p)^n for the probability of zero events in the binomial distribution\". The original treatment is Hanley and Lippman-Hand, \"If nothing goes wrong, is everything all right? Interpreting zero numerators\", JAMA 1983;249(13):1743-1745, PMID 6827763 - a title that is also the best one-line summary of the problem.\n\nTranslated into the sample sizes research actually uses:\n\n| Interviews with zero mentions | 95% upper bound on the true rate |\n| --- | --- |\n| 10 | about 30% |\n| 20 | about 15% |\n| 30 | 10% |\n| 60 | 5% |\n| 100 | 3% |\n| 300 | 1% |\n\nRead the first row again. Ten interviews in which nobody mentioned a problem are consistent with that problem affecting nearly a third of your users.\n\nOne honesty note on the arithmetic: the rule is stated for n greater than 30, and below that it is mildly conservative rather than wrong. At n = 20 the exact binomial upper bound is about 14 percent where 3/n gives 15 percent. The approximation errs toward caution, which is the direction you want.\n\n## The import, stated honestly\n\nThree adjacent ideas get confused with this one, and the distinctions are what make the rule usable.\n\n**Data saturation** tells you when to stop collecting because new themes have stopped appearing. It is a stopping rule about the themes you *are* finding. It is silent on bounds, and it will happily tell you to stop at 12 interviews - a sample that bounds nothing below 25 percent. See [Data Saturation in Qualitative Research](/docs/data-saturation-qualitative-research).\n\n**Capture-recapture estimation** infers how many themes you missed from the overlap between two independent analysts or passes. It is powerful and it has a hard requirement: it needs observations to work with. It cannot speak about a theme that appeared zero times to everybody. See [Capture-Recapture for Research](/docs/capture-recapture-theme-coverage). The rule of three is the tool for exactly the cell capture-recapture cannot reach: the zero cell.\n\n**Detection limits** describe the smallest effect your method can reliably see at all. That is a property of the instrument; the rule of three is a statement about the population given a clean null result. See [The Theme Is Real and the Percentage Is Noise](/docs/theme-detection-limit-vs-quantification).\n\nSo the claim here is narrow and specific: **when a thing did not happen, the rule of three gives you the only defensible number you are entitled to report.**\n\n## Why you can prove common but not rare\n\nThere is an asymmetry in qualitative evidence that no amount of rigour removes.\n\nIf 14 of 20 participants describe the same friction, you have strong evidence it is widespread. Small samples are genuinely good at detecting common things - that is the entire basis of small-n qualitative research and it is sound.\n\nIf 0 of 20 participants describe a friction, you have weak evidence about anything. You have excluded \"affects most users\" and you have not excluded \"affects one user in eight\". The evidence is asymmetric because the information content of a zero is low: many different underlying rates all produce zero in a small sample without straining credulity.\n\nThe practical failure follows directly. Teams run a study, find a clear dominant theme, ship against it, and treat everything that did not surface as settled. The dominant theme was well measured. The silence was not measured at all - and in interviews, silence has its own generating mechanisms, including the one described in [Why the Loudest Complaint Hides the Real One](/docs/dominant-complaint-masking-interviews).\n\n## The sample size is set by the consequence\n\nOnce you accept that an absence is a bound, sample size stops being a methodological question and becomes a risk question: how small a rate do you need to rule out before you are willing to act?\n\n- Deciding not to build a nice-to-have? A 15 percent bound is fine. Twenty interviews.\n- Deciding that a data-loss bug reported by one enterprise account is not systemic? You need a 1 percent bound. That is 300 observations, and no qualitative study is going to get you there - this is a telemetry question wearing a research costume.\n- Deciding that a compliance or accessibility failure does not affect your users? The tolerable rate is near zero, so the required n is enormous, and the honest answer is that interviews cannot clear this and you should audit instead of ask.\n\nThis inverts the usual planning conversation, and it is the calculation Koji users tend to run first. The question is not \"how many interviews can we afford?\" but \"what rate must we rule out, and does that number exceed what interviewing can deliver?\" Often it does, and knowing that early is worth more than the study. [Expected Value of Information](/docs/value-of-information-research-decisions) covers the companion decision of whether to run the study at all, and [How Many Interviews Are Enough?](/docs/how-many-interviews-enough) covers sizing for discovery rather than for bounds.\n\n## What Nielsen's curve does and does not tell you\n\nJakob Nielsen's problem-discovery curve, published in March 2000, is the most cited sample-size argument in UX. It models the proportion of usability problems found as a function of users tested, where \"the typical value of L is 31%, averaged across a large number of projects we studied\" - L being the share of problems a single user reveals.\n\nThat curve answers a discovery question: how much of the findable problem set will I have seen? It does not answer a bounding question: given that a specific problem did not appear, how common can it still be? These are different questions and the curve is regularly pressed into service for the second one, which it cannot answer.\n\nNielsen does state the boundary case with admirable bluntness - \"the most striking truth of the curve is that zero users give zero insights\" - and the rule of three is the quantitative continuation of that sentence. Zero users give zero insight; five users give you a 60 percent bound; twenty give you 15 percent. The curve tells you what you found. The rule of three tells you what you are still exposed to.\n\n## A protocol for reporting absences\n\n1. **Never write \"no participants reported X\" without the bound.** Write \"no participants reported X (n = 24, so the rate is below roughly 13 percent at 95 percent confidence)\". It is one clause and it prevents a whole class of bad decisions.\n2. **Pre-register the absences you care about.** Decide before fielding which specific risks you want to bound, because a bound is only valid for something you actually asked everyone about. An issue nobody was asked about has no n.\n3. **Ask, do not wait.** A theme that could have emerged but did not has an ambiguous n. A yes_no question put to all 40 participants has n = 40. Only the second supports a bound, which is why Koji puts every structured question to every participant.\n4. **Set the required bound before choosing the sample size**, driven by the cost of being wrong.\n5. **Escalate to telemetry when the required n exceeds a few hundred.** That is not a failure of the research; it is the research correctly telling you it is the wrong instrument.\n\n## How Koji handles this\n\nThe binding constraint on the rule of three is that a bound requires a denominator, and a denominator requires that everybody was actually asked. Traditional qualitative research struggles here: in an unstructured interview, whether a topic came up depends on the moderator, the time remaining and the participant's talkativeness, so there is no clean n to divide 3 by.\n\n- **Structured questions create real denominators.** Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and every one of them is put to every participant. A yes_no question asked of 60 people gives you a defensible n = 60 and therefore a defensible 5 percent bound. Emergent themes cannot do this, which is why absence claims from unstructured research are usually unquantifiable.\n- **Sample sizes that make bounds affordable.** A 1 percent bound needs 300 observations. Scheduling 300 moderated interviews is not realistic; running 300 AI-moderated interviews is, and this is the clearest case where the economics of AI-native research change which questions are answerable rather than merely making an old workflow faster.\n- **Consistent probing protects the denominator.** Koji's AI moderator asks the planned questions of everyone rather than running out of time on participant 14, so your n is the number of completed interviews instead of an unknown subset.\n- **Automatic thematic analysis and real-time reporting** let you see which pre-registered risks are still at zero while fielding is in progress, so you can extend a wave to reach the bound you need instead of discovering the shortfall at analysis time.\n- **Voice interviews and customisable AI consultants** keep the instrument identical across a large wave, which is what makes the independence assumption behind the rule reasonable.\n\nThe reframe is the valuable part, and it does not require a statistics background: stop reporting zero, start reporting 3/n.\n\n## Frequently asked questions\n\n### What is the rule of three in statistics?\n\nIt states that if you observe zero events in n independent observations, the 95 percent confidence interval for the true rate runs from 0 to 3/n. With 30 observations and no events, the rate could still be as high as 10 percent. It follows from the Poisson approximation to the binomial, and was set out by Hanley and Lippman-Hand in JAMA in 1983.\n\n### If nobody mentioned an issue in 20 interviews, is it safe to deprioritise?\n\nIt is safe to conclude the issue probably affects fewer than about 15 percent of users, and nothing stronger. Whether that justifies deprioritising depends entirely on the cost of being wrong. For a cosmetic improvement, yes. For data loss, a security concern or an accessibility failure, a 15 percent bound is nowhere near sufficient.\n\n### How many interviews do I need to rule out a 1% problem?\n\nAbout 300, since 3 divided by 300 is 0.01. That is the honest answer, and it explains why interviews are the wrong instrument for rare-event questions unless you can run them at scale. Koji makes a 300-participant wave practical, but for genuinely rare technical failures telemetry remains the better tool.\n\n### Is this the same as data saturation?\n\nNo. Saturation is a stopping rule based on new themes ceasing to appear, and it speaks only about themes you are finding. The rule of three speaks about a theme that never appeared, and converts that absence into an upper bound. A study can be fully saturated and still bound nothing below 25 percent.\n\n### Does the rule of three work for small samples?\n\nIt is stated for n greater than 30. Below that it remains usable and is mildly conservative: at n = 20 the exact binomial upper bound is about 14 percent while 3/n gives 15 percent. Since it errs toward a wider bound, using it on small samples will not make you overconfident.\n\n### How is this different from capture-recapture estimation?\n\nCapture-recapture estimates what you missed from the overlap between two independent passes, so it needs at least some observations to work from. The rule of three handles the case capture-recapture cannot reach: a theme that appeared zero times for everyone. The two are complementary rather than competing.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, and why asked-of-everyone is what creates a denominator\n- [Capture-Recapture for Research](/docs/capture-recapture-theme-coverage) - estimating what you missed when you did observe something\n- [Data Saturation in Qualitative Research](/docs/data-saturation-qualitative-research) - a stopping rule, not a bound\n- [How Many Interviews Are Enough?](/docs/how-many-interviews-enough) - sizing for discovery rather than for exclusion\n- [The Theme Is Real and the Percentage Is Noise](/docs/theme-detection-limit-vs-quantification) - what your method can detect at all\n- [Why the Loudest Complaint Hides the Real One](/docs/dominant-complaint-masking-interviews) - why a silence may be manufactured rather than real","category":"Research Methods","lastModified":"2026-09-27T03:31:41.873817+00:00","metaTitle":"The Rule of Three: What Zero Mentions in Your Research Actually Rules Out (2026)","metaDescription":"Zero mentions is not a rate of zero. The rule of three turns an absence into an upper bound of 3/n, the only number you can defend.","keywords":["zero numerator research","rule of three sample size","nobody mentioned it in interviews","absence of evidence research","upper bound zero occurrences","how many interviews to rule out"],"aiSummary":"Observing zero occurrences in n interviews does not establish a rate of zero; it establishes a 95 percent upper bound of approximately 3/n. Twenty interviews with no mentions are consistent with a 15 percent prevalence. The rule originates with Hanley and Lippman-Hand (JAMA 1983) and is operationalised by Eypasch et al (BMJ 1995) as maximum risk equals 3/n for n above 30. This makes sample size a function of the rate you must rule out rather than of saturation, and it requires a real denominator, which only questions asked of every participant can supply.","aiPrerequisites":["Familiarity with reading a confidence interval","Basic understanding of qualitative sample sizes and saturation"],"aiLearningOutcomes":["Convert an absence in your data into a defensible upper bound","Explain why small samples detect common issues but cannot establish rarity","Choose a sample size from the rate you must rule out and the cost of being wrong","Distinguish the rule of three from saturation, capture-recapture and detection limits","Recognise when a question requires telemetry rather than interviews"],"aiDifficulty":"intermediate","aiEstimatedTime":"9 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}