Back to docs
Research Methods

Nobody Mentioned It: What Your Sample Could Not Have Detected (2026)

Silence in your interviews is not evidence a problem is rare. Here is how to compute the prevalence floor your study was blind to, and why that number belongs in every report.

Answer first: If you ran 10 interviews and nobody raised an issue, you have not shown the issue is rare. You have shown it is probably not affecting more than about 20 percent of users. A problem affecting 5 percent of your base had a 60 percent chance of going completely unmentioned in those 10 conversations, and in a product with 100,000 users, 5 percent is 5,000 people. The number worth reporting is not how many interviews you ran. It is the prevalence floor below which your study was blind.

The question nobody asks after a study

Research reports are built around what was found. Themes, quotes, frequencies, recommendations. The corresponding question about what could not have been found is almost never on the page, even though it is the one that determines what the silence means.

Environmental sampling treats this as the central design question rather than an afterthought, because the stakes make it unavoidable. If you are looking for a patch of contamination in a field, you must decide in advance how small a patch you are willing to miss, because that decision sets how many samples you take. Pacific Northwest National Laboratory describes the design goal for this class of survey as being "to delineate regions of high concentration levels, or 'hot spots', while also reducing the number of laboratory analyses required and improving the estimates of the mean concentration through the use of multiple increment sampling." Their software computes, for a given number of samples, the probability of detecting contamination above a specified level, or inversely the number of samples needed to reach a desired detection power.

Notice what that framing does. It refuses to report a clean result without also reporting the size of the thing that could have hidden. Customer research almost never does this, and the arithmetic is not hard. Whether you run studies in Koji or anywhere else, the calculation below takes a single line.

The arithmetic

If an issue affects a proportion p of your population, and you sample n people independently, the probability that at least one of them raises it is:

P(detect) = 1 - (1 - p)^n

That is elementary probability: the chance of missing it with one person is (1 - p), missing it with all n independently is (1 - p)^n, and detection is everything else.

For a study of 10 interviews:

True prevalenceChance you hear it at least onceChance you miss it entirely
50 percent99.9 percent0.1 percent
20 percent89.3 percent10.7 percent
10 percent65.1 percent34.9 percent
5 percent40.1 percent59.9 percent
2 percent18.3 percent81.7 percent
1 percent9.6 percent90.4 percent

Read the bottom half of that table slowly. At 10 interviews, a problem affecting one user in twenty is more likely to be missed than found. A problem affecting one in a hundred will be missed nine times out of ten. Neither of those is a small problem in absolute terms; at 100,000 users they are 5,000 and 1,000 people respectively.

Turning it around: the sample you would need

The same formula run backwards gives the sample size for a chosen detection confidence. For a 90 percent chance of hearing an issue at least once:

Prevalence you want to catchInterviews required
10 percent22
5 percent45
2 percent114
1 percent230

These numbers explain something that otherwise looks like a contradiction in research practice. The familiar advice that a handful of participants is enough is sound for its purpose: finding the high-prevalence usability problems that affect most users. At 50 percent prevalence, five participants give you a 96.9 percent chance of detection, so five is genuinely plenty. The same five participants give you a 22.6 percent chance of catching a 5 percent issue. The sample size is not right or wrong in itself. It is right or wrong relative to the prevalence you need to see.

This is the complement to How Many Interviews Are Enough?, which works the forward question of sizing a study. This article works the inverse: given the study you already ran, what were you blind to?

Why this is the honest version of a negative finding

There is a specific and common reporting failure this corrects.

A team runs 12 interviews about a new onboarding flow. Nobody mentions the data import step. The report says onboarding is working well, and the summary slide says "no concerns raised about import."

Both sentences are true as descriptions of the transcripts. Neither supports the conclusion the room will draw, which is that import is fine. With 12 interviews, an import problem affecting 10 percent of users had a 28.2 percent chance of going unmentioned, and one affecting 5 percent had a 54 percent chance. The absence of the theme is weak evidence at best, and the report presented it as strong evidence by not quantifying it. A Koji report gives you the participant count and the per-theme frequencies the calculation needs, so the floor is one line of arithmetic away.

The fix is one sentence in the report: "This study could reliably detect issues affecting more than roughly 20 percent of users. Below that, absence of a theme is not evidence of absence." That sentence costs nothing, and it is the difference between a finding and an unwarranted reassurance.

What silence actually licenses

To be precise about what you can and cannot say after a null result:

  • Supported: the issue is probably not extremely common. If it affected half your users you would almost certainly have heard it.
  • Supported: the issue is not among the most salient concerns for this population, which is a real and useful finding about priority.
  • Not supported: the issue is rare.
  • Not supported: the issue does not exist.
  • Not supported: the issue is not worth fixing, if severity is high. A low-prevalence, high-severity issue is exactly the combination this design is worst at finding and that Why Complaint Counts Cannot Become Rates warns against ranking by frequency.

The independence assumption, stated honestly

The formula assumes each participant is an independent draw at prevalence p. Real studies violate this in both directions and it is worth knowing which way.

Violations that make you more blind than the table says. If your participants come from one source, one cohort, one region or one recruitment channel, they are correlated, and correlated draws cover less of the population than independent ones. Your effective n is smaller than your actual n. This is why Survivorship Bias in Customer Research compounds the problem: a sample drawn only from active users cannot detect an issue that causes people to stop being active.

Violations that make you less blind. If your screener deliberately targets the subpopulation most likely to have the issue, prevalence within your sample is higher than p in the general population, and detection is correspondingly better. This is a legitimate and underused design: to look for a rare problem, do not increase n across the whole base, increase p by sampling where the problem lives.

Treat the table as the calculation for a clean, broad, independent sample, and adjust your interpretation for whichever of these applies. Estimating the themes you missed from the data itself rather than from a formula is a different and complementary technique, covered in Capture-Recapture for Research.

Common mistakes

  • Reporting the sample size instead of the detection floor. "n = 12" tells a stakeholder nothing about what the study could see. "Blind below 20 percent" tells them everything.
  • Treating a null result as a green light. Absence of a theme at small n is close to uninformative for low-prevalence issues.
  • Adding interviews to chase a rare problem. Going from 10 to 20 interviews takes your 5 percent detection from 40.1 percent to 64.2 percent. Targeting your screener is far more efficient than doubling n.
  • Ignoring severity. The detection floor tells you about prevalence only. A 1 percent issue that loses enterprise accounts outranks a 40 percent cosmetic annoyance.
  • Applying the formula to a correlated sample without saying so. If everyone came from the same channel, your effective sample is smaller than n.

How Koji changes the economics of the floor

The detection floor is arithmetic and no tool alters it. What a tool alters is where the floor sits for a realistic budget, and that is decided almost entirely by the cost of one additional conversation.

Under a human-moderated model, each interview costs an hour of scheduling, an hour of moderation and a further block of transcription and synthesis. At that price, 45 interviews to reach a 90 percent chance of catching a 5 percent issue is a quarter-long project, so teams run 10, and the floor sits near 20 percent whether or not anyone computes it.

Koji changes the shape of that constraint. Interviews run asynchronously and in parallel with no moderator, so participants are limited by recruitment rather than by calendar. A study of 45 finishes in the time a study of 10 takes, and the analysis arrives with it rather than a week later. A 5 percent detection floor becomes an ordinary study rather than a special project.

Several capabilities matter specifically for finding low-prevalence issues:

  • AI follow-up questions raise effective detection. Prevalence in the formula is the chance a participant raises the issue, not the chance they have it. A participant who has a problem but does not think to mention it is a miss. Because Koji's AI interviewer probes each answer rather than accepting the first response, the gap between having and mentioning narrows, which raises p for the same underlying population.
  • Structured questions give you a direct denominator. With six question types available (open_ended, scale, single_choice, multiple_choice, ranking, yes_no), you can ask about a suspected issue explicitly with yes_no or single_choice rather than waiting for it to surface spontaneously. Asking directly converts a detection problem into a measurement problem, which needs far less sample.
  • Targeted screening is cheap. Because studies are inexpensive to run, sampling where prevalence is high, the cohort that churned, the plan tier that complains, is practical rather than a luxury.
  • Real-time reports let you extend a study that is trending toward a null. If 10 interviews in you have heard nothing, you can add participants while the study is live instead of writing up an underpowered null.

Koji cannot make silence mean more than it does. It can make the floor low enough that silence is worth something.

Frequently asked questions

What detection floor should I aim for?

Work backwards from consequence rather than picking a number. If a 5 percent issue would be expensive at your scale, design to detect 5 percent, which is about 45 interviews for 90 percent confidence. If you only need to catch problems that affect most users, five to ten participants genuinely suffices. The floor should be set by what it costs you to miss something.

Does this contradict the advice to test with five users?

No, and the two fit together precisely. The five-user guidance targets high-prevalence usability problems, and for those it is well founded: at 50 percent prevalence, five participants detect an issue 96.9 percent of the time. The guidance was never a claim about rare problems, and the arithmetic shows why it cannot be.

Should the detection floor go in the report?

Yes, and it is the single cheapest improvement available to most research reports. One sentence stating the prevalence below which absence of a theme is uninformative prevents the most common misreading of qualitative results, which is treating silence as reassurance.

How does this interact with saturation?

They answer different questions and are often confused. Saturation asks whether new interviews are still producing new themes, which is a property of your sample. The detection floor asks what prevalence of issue your sample size could have caught, which is a property of the population. A study can reach apparent saturation and still be blind to a 5 percent problem, because rare themes stop appearing long before common ones do.

What if I cannot recruit enough participants?

Then raise prevalence instead of sample size, which is usually both cheaper and faster. Screen for the cohort most likely to have the issue rather than sampling the whole base. Ten interviews inside a population where the issue runs at 40 percent detect it 99.4 percent of the time, which beats 45 interviews across a base where it runs at 5 percent.

Can I use this for themes I did not anticipate?

Only loosely, and the limitation is worth being clear about. The formula requires a specific p for a specific issue, so it applies cleanly when you are checking whether a particular problem exists. For unanticipated themes there is no p to plug in, and the better tool is a coverage estimate from the data itself, such as capture-recapture across independent coders.

Related Resources