{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-30T18:52:54.625Z"},"content":[{"type":"documentation","id":"1429ddf6-02dd-4113-b7b1-0c6bf305fd4b","slug":"detection-floor-nobody-mentioned-interviews","title":"Nobody Mentioned It: What Your Sample Could Not Have Detected (2026)","url":"https://www.koji.so/docs/detection-floor-nobody-mentioned-interviews","summary":"Silence in a small qualitative study is weak evidence. P(detect) = 1 - (1-p)^n: at n=10, a 5 percent issue is missed 59.9 percent of the time and a 1 percent issue 90.4 percent. Reaching 90 percent detection needs 22 interviews at 10 percent prevalence, 45 at 5 percent, 114 at 2 percent, 230 at 1 percent. Reports should state the prevalence floor below which absence of a theme is uninformative.","content":"**Answer first:** If you ran 10 interviews and nobody raised an issue, you have not shown the issue is rare. You have shown it is probably not affecting more than about 20 percent of users. A problem affecting 5 percent of your base had a 60 percent chance of going completely unmentioned in those 10 conversations, and in a product with 100,000 users, 5 percent is 5,000 people. The number worth reporting is not how many interviews you ran. It is the prevalence floor below which your study was blind.\n\n## The question nobody asks after a study\n\nResearch reports are built around what was found. Themes, quotes, frequencies, recommendations. The corresponding question about what *could not* have been found is almost never on the page, even though it is the one that determines what the silence means.\n\nEnvironmental sampling treats this as the central design question rather than an afterthought, because the stakes make it unavoidable. If you are looking for a patch of contamination in a field, you must decide in advance how small a patch you are willing to miss, because that decision sets how many samples you take. Pacific Northwest National Laboratory describes the design goal for this class of survey as being \"to delineate regions of high concentration levels, or 'hot spots', while also reducing the number of laboratory analyses required and improving the estimates of the mean concentration through the use of multiple increment sampling.\" Their software computes, for a given number of samples, the probability of detecting contamination above a specified level, or inversely the number of samples needed to reach a desired detection power.\n\nNotice what that framing does. It refuses to report a clean result without also reporting the size of the thing that could have hidden. Customer research almost never does this, and the arithmetic is not hard. Whether you run studies in Koji or anywhere else, the calculation below takes a single line.\n\n## The arithmetic\n\nIf an issue affects a proportion p of your population, and you sample n people independently, the probability that at least one of them raises it is:\n\n**P(detect) = 1 - (1 - p)^n**\n\nThat is elementary probability: the chance of missing it with one person is (1 - p), missing it with all n independently is (1 - p)^n, and detection is everything else.\n\nFor a study of 10 interviews:\n\n| True prevalence | Chance you hear it at least once | Chance you miss it entirely |\n| --- | --- | --- |\n| 50 percent | 99.9 percent | 0.1 percent |\n| 20 percent | 89.3 percent | 10.7 percent |\n| 10 percent | 65.1 percent | 34.9 percent |\n| 5 percent | 40.1 percent | 59.9 percent |\n| 2 percent | 18.3 percent | 81.7 percent |\n| 1 percent | 9.6 percent | 90.4 percent |\n\nRead the bottom half of that table slowly. **At 10 interviews, a problem affecting one user in twenty is more likely to be missed than found.** A problem affecting one in a hundred will be missed nine times out of ten. Neither of those is a small problem in absolute terms; at 100,000 users they are 5,000 and 1,000 people respectively.\n\n### Turning it around: the sample you would need\n\nThe same formula run backwards gives the sample size for a chosen detection confidence. For a 90 percent chance of hearing an issue at least once:\n\n| Prevalence you want to catch | Interviews required |\n| --- | --- |\n| 10 percent | 22 |\n| 5 percent | 45 |\n| 2 percent | 114 |\n| 1 percent | 230 |\n\nThese numbers explain something that otherwise looks like a contradiction in research practice. The familiar advice that a handful of participants is enough is sound *for its purpose*: finding the high-prevalence usability problems that affect most users. At 50 percent prevalence, five participants give you a 96.9 percent chance of detection, so five is genuinely plenty. The same five participants give you a 22.6 percent chance of catching a 5 percent issue. **The sample size is not right or wrong in itself. It is right or wrong relative to the prevalence you need to see.**\n\nThis is the complement to [How Many Interviews Are Enough?](/docs/how-many-interviews-enough), which works the forward question of sizing a study. This article works the inverse: given the study you already ran, what were you blind to?\n\n## Why this is the honest version of a negative finding\n\nThere is a specific and common reporting failure this corrects.\n\nA team runs 12 interviews about a new onboarding flow. Nobody mentions the data import step. The report says onboarding is working well, and the summary slide says \"no concerns raised about import.\"\n\nBoth sentences are true as descriptions of the transcripts. Neither supports the conclusion the room will draw, which is that import is fine. With 12 interviews, an import problem affecting 10 percent of users had a 28.2 percent chance of going unmentioned, and one affecting 5 percent had a 54 percent chance. The absence of the theme is weak evidence at best, and the report presented it as strong evidence by not quantifying it. A Koji report gives you the participant count and the per-theme frequencies the calculation needs, so the floor is one line of arithmetic away.\n\nThe fix is one sentence in the report: **\"This study could reliably detect issues affecting more than roughly 20 percent of users. Below that, absence of a theme is not evidence of absence.\"** That sentence costs nothing, and it is the difference between a finding and an unwarranted reassurance.\n\n### What silence actually licenses\n\nTo be precise about what you can and cannot say after a null result:\n\n- **Supported:** the issue is probably not extremely common. If it affected half your users you would almost certainly have heard it.\n- **Supported:** the issue is not among the most salient concerns for this population, which is a real and useful finding about priority.\n- **Not supported:** the issue is rare.\n- **Not supported:** the issue does not exist.\n- **Not supported:** the issue is not worth fixing, if severity is high. A low-prevalence, high-severity issue is exactly the combination this design is worst at finding and that [Why Complaint Counts Cannot Become Rates](/docs/complaint-counts-cannot-be-rates) warns against ranking by frequency.\n\n## The independence assumption, stated honestly\n\nThe formula assumes each participant is an independent draw at prevalence p. Real studies violate this in both directions and it is worth knowing which way.\n\n**Violations that make you more blind than the table says.** If your participants come from one source, one cohort, one region or one recruitment channel, they are correlated, and correlated draws cover less of the population than independent ones. Your effective n is smaller than your actual n. This is why [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) compounds the problem: a sample drawn only from active users cannot detect an issue that causes people to stop being active.\n\n**Violations that make you less blind.** If your screener deliberately targets the subpopulation most likely to have the issue, prevalence within your sample is higher than p in the general population, and detection is correspondingly better. This is a legitimate and underused design: to look for a rare problem, do not increase n across the whole base, increase p by sampling where the problem lives.\n\nTreat the table as the calculation for a clean, broad, independent sample, and adjust your interpretation for whichever of these applies. Estimating the themes you missed from the data itself rather than from a formula is a different and complementary technique, covered in [Capture-Recapture for Research](/docs/capture-recapture-theme-coverage).\n\n## Common mistakes\n\n- **Reporting the sample size instead of the detection floor.** \"n = 12\" tells a stakeholder nothing about what the study could see. \"Blind below 20 percent\" tells them everything.\n- **Treating a null result as a green light.** Absence of a theme at small n is close to uninformative for low-prevalence issues.\n- **Adding interviews to chase a rare problem.** Going from 10 to 20 interviews takes your 5 percent detection from 40.1 percent to 64.2 percent. Targeting your screener is far more efficient than doubling n.\n- **Ignoring severity.** The detection floor tells you about prevalence only. A 1 percent issue that loses enterprise accounts outranks a 40 percent cosmetic annoyance.\n- **Applying the formula to a correlated sample without saying so.** If everyone came from the same channel, your effective sample is smaller than n.\n\n## How Koji changes the economics of the floor\n\nThe detection floor is arithmetic and no tool alters it. What a tool alters is where the floor sits for a realistic budget, and that is decided almost entirely by the cost of one additional conversation.\n\nUnder a human-moderated model, each interview costs an hour of scheduling, an hour of moderation and a further block of transcription and synthesis. At that price, 45 interviews to reach a 90 percent chance of catching a 5 percent issue is a quarter-long project, so teams run 10, and the floor sits near 20 percent whether or not anyone computes it.\n\nKoji changes the shape of that constraint. Interviews run asynchronously and in parallel with no moderator, so participants are limited by recruitment rather than by calendar. A study of 45 finishes in the time a study of 10 takes, and the analysis arrives with it rather than a week later. A 5 percent detection floor becomes an ordinary study rather than a special project.\n\nSeveral capabilities matter specifically for finding low-prevalence issues:\n\n- **AI follow-up questions raise effective detection.** Prevalence in the formula is the chance a participant *raises* the issue, not the chance they *have* it. A participant who has a problem but does not think to mention it is a miss. Because Koji's AI interviewer probes each answer rather than accepting the first response, the gap between having and mentioning narrows, which raises p for the same underlying population.\n- **Structured questions give you a direct denominator.** With six question types available (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), you can ask about a suspected issue explicitly with yes_no or single_choice rather than waiting for it to surface spontaneously. Asking directly converts a detection problem into a measurement problem, which needs far less sample.\n- **Targeted screening is cheap.** Because studies are inexpensive to run, sampling where prevalence is high, the cohort that churned, the plan tier that complains, is practical rather than a luxury.\n- **Real-time reports let you extend a study that is trending toward a null.** If 10 interviews in you have heard nothing, you can add participants while the study is live instead of writing up an underpowered null.\n\nKoji cannot make silence mean more than it does. It can make the floor low enough that silence is worth something.\n\n## Frequently asked questions\n\n### What detection floor should I aim for?\n\nWork backwards from consequence rather than picking a number. If a 5 percent issue would be expensive at your scale, design to detect 5 percent, which is about 45 interviews for 90 percent confidence. If you only need to catch problems that affect most users, five to ten participants genuinely suffices. The floor should be set by what it costs you to miss something.\n\n### Does this contradict the advice to test with five users?\n\nNo, and the two fit together precisely. The five-user guidance targets high-prevalence usability problems, and for those it is well founded: at 50 percent prevalence, five participants detect an issue 96.9 percent of the time. The guidance was never a claim about rare problems, and the arithmetic shows why it cannot be.\n\n### Should the detection floor go in the report?\n\nYes, and it is the single cheapest improvement available to most research reports. One sentence stating the prevalence below which absence of a theme is uninformative prevents the most common misreading of qualitative results, which is treating silence as reassurance.\n\n### How does this interact with saturation?\n\nThey answer different questions and are often confused. Saturation asks whether new interviews are still producing new themes, which is a property of your sample. The detection floor asks what prevalence of issue your sample size could have caught, which is a property of the population. A study can reach apparent saturation and still be blind to a 5 percent problem, because rare themes stop appearing long before common ones do.\n\n### What if I cannot recruit enough participants?\n\nThen raise prevalence instead of sample size, which is usually both cheaper and faster. Screen for the cohort most likely to have the issue rather than sampling the whole base. Ten interviews inside a population where the issue runs at 40 percent detect it 99.4 percent of the time, which beats 45 interviews across a base where it runs at 5 percent.\n\n### Can I use this for themes I did not anticipate?\n\nOnly loosely, and the limitation is worth being clear about. The formula requires a specific p for a specific issue, so it applies cleanly when you are checking whether a particular problem exists. For unanticipated themes there is no p to plug in, and the better tool is a coverage estimate from the data itself, such as capture-recapture across independent coders.\n\n## Related Resources\n\n- [How Many Interviews Are Enough?](/docs/how-many-interviews-enough) - the forward version of this calculation\n- [Structured Questions Guide](/docs/structured-questions-guide) - asking directly with the six question types instead of waiting for a theme to surface\n- [Capture-Recapture for Research](/docs/capture-recapture-theme-coverage) - estimating missed themes from the data rather than a formula\n- [Why Complaint Counts Cannot Become Rates](/docs/complaint-counts-cannot-be-rates) - why frequency is not the severity axis\n- [Survivorship Bias in Customer Research](/docs/survivorship-bias-customer-research) - the sampling frame problem that shrinks your effective n\n- [The Composite Sample Problem](/docs/composite-sample-aggregate-feedback-variance) - the aggregation side of the same detection question\n","category":"Research Methods","lastModified":"2026-09-29T03:36:26.140063+00:00","metaTitle":"Nobody Mentioned It: What Your Sample Could Not Have Detected (2026)","metaDescription":"Nobody mentioning a problem is not evidence it is rare. Compute the prevalence floor your sample could not have detected, and report it.","keywords":["detection probability interviews","prevalence floor research","absence of evidence user research","how many interviews to find rare problem","null result qualitative research"],"aiSummary":"Silence in a small qualitative study is weak evidence. P(detect) = 1 - (1-p)^n: at n=10, a 5 percent issue is missed 59.9 percent of the time and a 1 percent issue 90.4 percent. Reaching 90 percent detection needs 22 interviews at 10 percent prevalence, 45 at 5 percent, 114 at 2 percent, 230 at 1 percent. Reports should state the prevalence floor below which absence of a theme is uninformative.","aiPrerequisites":["Basic probability","Familiarity with qualitative sample sizing"],"aiLearningOutcomes":["Compute the detection probability for a given prevalence and sample size","Derive the sample size needed to detect a target prevalence","State a defensible detection floor in a research report","Distinguish saturation from detection power"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}