The Theme Is Real and the Percentage Is Noise: Detection vs Quantification Limits (2026)
Finding a theme and measuring how common it is are two different instruments with two different thresholds. Analytical chemistry separates them formally, and qualitative research should too.
Two of your twenty interview participants raised the same complaint. You are entitled to say the problem is real. You are not entitled to say it affects 10 percent of your users. Those are two different claims, they require different amounts of evidence, and the gap between them is larger than almost any research report admits.
Analytical chemistry formalised this distinction decades ago and gave the two thresholds separate names. Qualitative research collapses them into one number and ships it.
The short answer
- Detection and quantification are separate thresholds. You can reliably establish that something is present well before you can reliably say how much of it there is.
- In chemistry the limit of detection is set at 3 standard deviations of the blank and the limit of quantification at 10, so quantifying needs roughly 3.3 times the signal that detecting needs.
- User research reproduces that ratio almost exactly: Nielsen Norman Group recommends 5 users to find problems and 20 to measure them, about 4 times as many.
- The practical rule: report small-count themes as present, without a percentage. A theme raised by 2 of 20 people has a 95 percent confidence interval of roughly 3 to 30 percent, which is not a measurement.
Two thresholds, not one
Chemists define the lower threshold as the point where presence becomes distinguishable from absence. The limit of detection is "the lowest quantity of a substance that can be distinguished from the absence of that substance." It is computed against a blank - a sample known to contain none of the target - and is conventionally placed at "3 x standard deviation of the blank."
The upper threshold answers a different question. The limit of quantification is "the lowest value of a signal (or concentration, activity, response...) that can be quantified," and it sits at "10 x standard deviation of the blank."
The ratio is the point. Ten over three is roughly 3.3, so it takes more than three times as much signal to state a magnitude as to state an existence. Between the two thresholds lies a legitimate and frequently useful zone: the substance is definitely there, and its concentration is not yet measurable.
There is a detail at the lower threshold that deserves to be better known. At the limit of detection, "the beta error (probability of a false negative) is 50%." The detection limit is not the point at which you always see the thing. It is the point at which you see it half the time. Only at the quantification limit is it true that "there is minimal chance of a false negative."
Apply that to a theme mentioned once in a study. A single mention is not evidence that the theme is rare. It is evidence that you are operating at your detection limit, where the coin-flip miss rate means an equally common theme in the next study may not appear at all.
User research independently landed on the same ratio
The striking thing is that usability research arrived at these two thresholds without borrowing the chemistry, and got a similar ratio.
On the detection side, Nielsen Norman Group reports that a qualitative study with five users is likely to surface about 85 percent of the usability problems present. Five participants is a remarkably good detector.
On the quantification side, Jakob Nielsen is explicit that this is a different job requiring a different sample: "I usually recommend testing with 20 users when collecting quantitative usability metrics." He states the cost ratio plainly: "You can usually run a qualitative study with 5 users, so quantitative studies are about 4 times as expensive."
Four times for measurement versus detection in usability. Three point three times in analytical chemistry. Two fields with no methodological contact converged on the same structural fact: measuring is three to four times more expensive in evidence than finding.
Raluca Budiu of Nielsen Norman Group also names why the two cannot substitute for each other: qualitative data give a direct assessment of a system usability, while quantitative data give only an indirect one. And a number does not carry its own explanation. Knowing that just 40 percent of participants managed to complete a task does not tell you why they struggled with it, or what to change so the next group does better.
What a small count actually supports
Here is the arithmetic that should sit next to every theme table. For a study of 20 participants, the Wilson 95 percent confidence interval on the proportion who mentioned a theme:
| Mentions | Point estimate | 95% interval | Interval width |
|---|---|---|---|
| 1 of 20 | 5% | 0.9% to 23.6% | 22.7 points |
| 2 of 20 | 10% | 2.8% to 30.1% | 27.3 points |
| 6 of 20 | 30% | 14.6% to 51.9% | 37.3 points |
| 10 of 20 | 50% | 29.9% to 70.1% | 40.2 points |
| 12 of 20 | 60% | 38.7% to 78.1% | 39.4 points |
The 2-of-20 row is the one that matters. The point estimate is 10 percent. The interval runs from 3 percent to 30 percent - a tenfold span. Reporting "10 percent of users hit this" is not a slightly imprecise statement, it is a statement whose plausible range covers everything from a rounding error to a top-three priority.
Two structural features of that table are worth noticing. First, the intervals are widest in the middle, not at the edges, so a 50 percent finding is your least precise proportion rather than your most solid one. Second, the 10-of-20 row is a check on the arithmetic: at exactly half the sample the interval is symmetric around 50 percent, which is what the Wilson formula must produce. If your own implementation does not return a symmetric interval at p equals one half, it is wrong.
What this is not
Three adjacent problems get confused with this one, and separating them keeps the fix clean.
It is not a shrinkage problem. When a small segment produces an implausible score, credibility weighting blends it toward your overall average to get a better estimate. That is the right tool when you need a number. Detection versus quantification is the prior question of whether you should be producing a number at all.
It is not a privacy problem. Minimum base sizes and k-anonymity suppress small cells because publishing them could identify individuals. That threshold is about disclosure risk and can be met by a cell that is still statistically meaningless, or breached by one that is perfectly precise.
It is not a coverage problem. Estimating how many themes you never saw is a separate question with its own estimators. This article is about themes you did see, and what kind of claim each one can carry.
The reporting discipline
The fix requires no new statistics, only two report sections instead of one.
Detected, not quantified. Themes above your detection threshold and below your quantification threshold. State them as present, with the raw count and a verbatim excerpt, and no percentage. Language that works: "raised independently by 2 of 20 participants" - a fact - rather than "affects 10 percent of users" - an unsupported estimate.
Quantified. Themes with enough observations to carry an interval, reported with the interval rather than the point estimate alone.
Three habits follow. Publish counts and denominators, never bare percentages, because "10 percent" hides whether it came from 2 of 20 or 200 of 2,000. Never rank a detected-only theme against a quantified one, since the ordering is an artefact of precision rather than prevalence. And never conclude that a theme is unimportant because it appeared once, because at the detection limit the false-negative rate is 50 percent.
The modern approach: raise the quantification limit, do not fake it
The traditional resolution of this problem is resignation. Interviews are expensive, so teams run 8 and then write percentages anyway because a stakeholder asked for one. The chemist equivalent would be reporting a concentration from a signal below the quantification limit, which would fail any audit.
The alternative is to move the threshold. If quantifying needs roughly 4 times the sample that detecting needs, then the constraint is the cost per interview, and that is exactly what AI-moderated research changes. Where a traditional study of 20 moderated interviews means 20 scheduled calls and weeks of transcription and coding, Koji runs AI-moderated interviews in parallel, so a sample large enough to quantify is no longer a budget decision made against a sample large enough to detect. You can have both.
Koji is built so the two claim types stay separable rather than blurring:
- Structured questions give you a genuine denominator. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and the five closed types are answered by every participant. That is the difference between a measurement and a tally of who happened to raise something unprompted. A scale item answered by all 40 participants is quantifiable; spontaneous mentions of the same topic in the open_ended responses are a detection channel. Running both in one Koji study gives you a detector and an instrument at the same time.
- Detection stays cheap and broad. Koji automatic thematic analysis reads every transcript rather than the subset a human had time for, so the open-ended side keeps its 5-user sensitivity across the whole sample instead of degrading as the study grows.
- Sample size stops being the binding constraint. Because Koji AI interviews, including voice interviews, run without scheduling, reaching 20 or 40 participants is a matter of distribution rather than calendar availability. Teams adopting AI-assisted research consistently report much faster time-to-insight, and the practical effect here is that the quantification threshold becomes reachable in a normal sprint.
- Real-time reporting shows the interval moving. As responses arrive you can watch a theme cross from detected into quantified, which is a far healthier way to decide a study is finished than a fixed target chosen in advance.
Democratising research does not mean pretending the thresholds are not there. It means making the larger sample cheap enough that you rarely have to report from the gap - and being precise when you do.
Frequently asked questions
What is the difference between detecting and quantifying a theme?
Detecting means establishing that a theme exists in your population. Quantifying means estimating what share of the population it affects. Detection needs far less evidence, so there is a real range in which a theme is confidently present and its prevalence is unmeasurable. Analytical chemistry marks the two points as the limit of detection and the limit of quantification.
Can I report a percentage from 8 interviews?
Report the count and the denominator instead. With 8 participants a single mention gives a 95 percent confidence interval spanning roughly 0 to 50 percent, so the percentage carries essentially no information while looking authoritative. Saying "2 of 8 participants raised this" is accurate and equally useful for prioritisation.
How many participants do I need to quantify a theme?
Nielsen Norman Group recommends about 20 for quantitative usability metrics against 5 for qualitative discovery, roughly a fourfold difference. Twenty still produces wide intervals, so treat it as the floor for quantification rather than a comfortable sample, and widen it if you intend to compare segments.
If a theme was mentioned only once, is it rare?
Not necessarily. At the detection limit the probability of a false negative is about 50 percent, so a theme that appeared once in this study could easily appear three times or zero times in an identical study. A single mention tells you the theme exists and tells you almost nothing about its frequency.
Is this the same as suppressing small segments for privacy?
No. Minimum base sizes and k-anonymity exist to prevent individuals being identified from small cells, which is a disclosure question. A cell can be large enough to publish safely and still far too small to quantify, so the two thresholds are independent and you need to satisfy both.
How does Koji help separate detection from quantification?
Koji pairs open-ended AI-moderated interviews, which are a sensitive detector analysed in full by automatic thematic analysis, with structured questions that every participant answers and which therefore support real measurement. Because Koji AI and voice interviews run in parallel without scheduling, reaching a sample large enough to quantify is practical, and real-time reporting lets you watch a theme cross from detected to quantified.
Related Resources
- Structured Questions in AI Interviews - the six question types, and which of them can carry a percentage
- Credibility Weighting for Small Segments - what to do once you have decided a number is required
- k-Anonymity for Segment Reporting - the separate, privacy-driven minimum base size
- Statistical Power and Minimum Detectable Effect - the same threshold logic applied to change rather than prevalence
- Singleton Themes and Unseen Coverage - what one-off comments tell you about what you missed
- Counting Themes by Mention or by Participant - choosing the denominator before you report one
- How Many Interviews Are Enough? - sample size for discovery versus measurement
Related Articles
When the Number Is Right and the Answer Is Wrong: Counting Themes by Mention or by Participant (2026)
One set of transcripts, two legitimate counting rules, two rankings with a Spearman correlation of -0.10. Both are correct. Only one answers your question.
Credibility Weighting: How Much of a Small Segment Score Should You Believe? (2026)
An eight-person segment scoring 4.6 against a 4.1 average is neither 4.6 nor unusable. Credibility weighting gives you the exact weight to apply, using a formula actuaries have relied on since 1918.
How Many Interviews Are Enough? A Guide to Sample Size
Understand saturation, practical guidelines, and research-backed recommendations for qualitative sample sizes.
k-Anonymity for Segment Reporting: How Small Is Too Small to Publish? (2026)
The rule for minimum base size in research reporting, stated exactly: every visible combination of attributes must be shared by at least k respondents - and why generalisation beats suppression.
Singleton Themes: Why One-Off Comments Are the Only Estimate You Have of What You Missed (2026)
Good-Turing says the chance the next respondent raises something new is the singleton count divided by total mentions. Every synthesis step deletes singletons first.
Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)
Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.