Singleton Themes: Why One-Off Comments Are the Only Estimate You Have of What You Missed (2026)
Good-Turing says the chance the next respondent raises something new is the singleton count divided by total mentions. Every synthesis step deletes singletons first.
Bottom line up front: The themes mentioned exactly once are the only evidence you have about the themes you never heard at all. Good-Turing's result is that the probability the next respondent says something belonging to no theme in your codebook is approximately f1/n -- the number of one-off themes divided by the number of coded mentions. In a typical corpus of 400 coded mentions across 47 themes with 12 one-offs, that is 3 percent: 97 percent of what people say is already covered, while only about 77 percent of the distinct kinds of things they might say have been found. Then the synthesis workflow runs. Merging near-duplicates, dropping "n=1 anecdotes", setting a minimum cluster size of three, and reporting the top five themes all remove one-off themes preferentially. Delete the one-offs and f1 becomes zero, which makes the estimated number of unseen themes exactly zero. The tidying step does not just lose a few stray quotes. It sets your estimate of what you are missing to zero, and it does so silently.
The one formula in this article
I. J. Good published the result in Biometrika in 1953, crediting the wartime work of Alan Turing. William Gale of AT&T Bell Laboratories states the usable form directly in Good-Turing Smoothing Without Tears (published as Gale and Sampson, Journal of Quantitative Linguistics 2, 1995): "A useful part of Good-Turing methodology is the estimate that the total probability of all unseen objects is N1/N."
In research terms:
- n = total coded mentions in the corpus (not interviews -- mentions).
- f1 = number of themes that appear exactly once. Singletons.
- f2 = number of themes that appear exactly twice. Doubletons.
- S-obs = number of distinct themes found.
- p0 = f1 / n = the probability that the next coded mention belongs to a theme you have never seen.
That is the whole thing. It needs one pass, one coder, and no second study. It is the single-sample counterpart to estimating coverage from the overlap between two passes, and unlike that method it costs nothing at all.
Two different coverages, and the gap between them
Run it on a realistic corpus: 25 interviews producing 400 coded mentions across 47 distinct themes, of which 12 appear exactly once and 5 appear exactly twice.
| Quantity | Value | Reads as |
|---|---|---|
| p0 = f1/n = 12/400 | 3.0 percent | Chance the next mention is a brand new theme |
| Mention coverage = 1 - p0 | 97.0 percent | Share of what gets said that your codebook already covers |
| Chao1 = S-obs + f1^2/(2 f2) = 47 + 144/10 | 61.4 themes | Estimated size of the theme space |
| Estimated unseen themes | about 14 | Distinct issues nobody in this study raised |
| Theme coverage = 47 / 61.4 | 76.5 percent | Share of the kinds of issue you have found |
The two coverage figures differ by 20.5 percentage points, and both are correct. Ninety-seven percent of the volume is covered because the common themes are common -- that is what makes them common. Meanwhile nearly a quarter of the distinct issues in the population have never been articulated to you.
Confusing these two numbers is the most consequential arithmetic error in qualitative synthesis. "We heard the same things over and over" is a statement about mention coverage. "We know what the problems are" is a claim about theme coverage. The first does not license the second.
The Chao1 estimator (Chao, 1984) is worth understanding for one reason, stated by Gotelli and Chao in the Encyclopedia of Biodiversity (2013): it is "based on the concept that rare species carry the most information about the number of undetected species", and it uses "only the numbers of singletons and doubletons (and the observed richness)". It is also, deliberately, a lower bound on the true richness. The real number of unseen themes is at least 14, not at most.
A two-second test for whether your codebook is closed
Gale gives a diagnostic that transfers perfectly and takes no tooling. Compare twice the doubletons to the singletons:
1* = 2 x f2 / f1
His rule: "The sign of a closed class is that 1* > 1." A closed class is one where you will soon have seen everything; an open class keeps producing new kinds no matter how long you sample. In the worked corpus, 1* = 2 x 5 / 12 = 0.83, which is below 1. The theme space is open. More interviewing will keep producing genuinely new issues, and the flat-looking curve in the report is about volume, not about kinds.
Run this on your last three studies. It costs one query and it will change how you write at least one of the conclusions.
Five routine steps that delete the measurement
Here is the part that makes this a structural problem rather than a statistical footnote. Every one of these is standard practice, defensible in isolation, and each removes singletons preferentially:
- The "Other" bucket. Themes too small to name get swept into a residual category, which is exactly f1 and f2 by construction.
- "That is an n=1, let us not over-index." Correct as a prioritisation instinct. Catastrophic as a data-retention rule, because n=1 is the definition of a singleton.
- Minimum cluster size. Automated clustering with a floor of three members removes f1 and f2 mechanically, before a human ever sees them.
- Merging near-duplicate themes. Reconciliation folds small themes into larger neighbours, converting singletons into increments on established counts.
- The top-five executive summary. Nothing is deleted from the data, but everything downstream reads a corpus in which f1 = 0.
Now the arithmetic of the tidy-up on the same corpus. Removing all singletons and doubletons deletes 22 of 400 mentions -- 5.5 percent of the corpus -- and 17 of 47 themes, which is 36.2 percent of the distinct kinds. Thirty themes remain. With f1 = 0 and f2 = 0, Chao1 returns exactly S-obs: 30 themes observed, zero estimated unseen, 100 percent theme coverage.
A study that genuinely covers 77 percent of the theme space now reports, by its own internal logic, that it covers all of it. Nobody falsified anything. The cleanup did it.
Why this is the failure mode that hides
Most research errors leave a trace. A biased sample shows up in the demographics table. A leading question shows up in the transcript. Publication bias shows up as a suspiciously tidy evidence base if anyone counts the studies that were never written up.
Singleton deletion leaves nothing behind. The remaining themes are all real, all well-evidenced, all correctly counted. The corpus looks better after the deletion by every quality metric a reviewer would apply: cleaner clusters, higher average theme frequency, better inter-coder agreement. The only casualty is the one statistic that estimates what is absent, and its absence is invisible because a missing estimate looks exactly like a confident one.
This is what makes it worth a policy rather than a habit: the number that measures your ignorance is stored in the data your process is designed to discard first.
Triage: which singletons are which
Not all one-offs deserve equal weight, and the answer is not to promote every stray comment to a theme. There are three kinds, and only the first two carry information about the unseen.
| Kind of singleton | How to recognise it | What it means |
|---|---|---|
| Rare in the population | Specific, coherent, articulated confidently, and the participant is a normal member of the sample | Genuine tail. It is the evidence that the theme space is open. |
| Hard to elicit | Appears late in an interview, after a probe, or only in voice sessions | Common in the population, rare in your instrument. Fix the guide, not the codebook. |
| Coding artefact | A splinter of an existing theme, or a label nobody else would apply the same way | Noise. Merge it, but record that you did, because merges reduce f1. |
The practical rule: you may merge singletons, but you may not merge them silently. Keep the pre-merge f1 and f2 alongside the post-merge codebook. That single discipline preserves the estimate through the entire tidy-up.
What to report instead
Add a coverage footer to every thematic report. Six numbers, one line each, computed before any merging:
- Coded mentions (n) and distinct themes (S-obs)
- Singletons (f1) and doubletons (f2)
- p0 = f1/n, stated as "probability the next respondent raises something new"
- Chao1 and the implied unseen count, labelled as a lower bound
- 1* = 2 f2 / f1, with "open" or "closed"
- Whether the figures are pre-merge or post-merge
Then write the conclusion in terms of the decision. Theme coverage of 77 percent is entirely adequate for prioritising the top of a roadmap and entirely inadequate for a claim that a segment's needs are understood. Same study, same number, opposite verdicts -- which is what a coverage estimate is for.
How Koji helps
The reason almost nobody reports f1 is mechanical: in a manual workflow, the singletons are already gone by the time anyone could count them. They were absorbed during affinity mapping on a whiteboard, or dropped in the spreadsheet consolidation, and the pre-merge counts were never written down. The fix has to live in the tooling.
- Every mention is retained and attributed. Koji's automatic thematic analysis keeps the full mention-level record with the interview it came from, so n, f1 and f2 are computable at any point -- including before a merge. Nothing has to be reconstructed from memory.
- AI-moderated interviews probe the second kind of singleton. A human moderator running to a schedule does not always chase an unexpected remark. An AI moderator has no time pressure and follows up consistently, which converts hard-to-elicit themes into properly evidenced ones instead of leaving them as one-offs.
- Voice interviews change what reaches the codebook at all. Themes that people will not type into a survey box, they will say out loud. Modality is a lever on f1 that question wording alone cannot reach.
- Structured questions bound the space so the tail is visible. Koji's six question types -- open_ended, scale, single_choice, multiple_choice, ranking, and yes_no -- fix the closed part of the instrument, which means the open_ended responses are the only place new kinds can appear, and the "Other" answers on single_choice and multiple_choice items are a clean, countable singleton pool rather than a black hole. See the structured questions guide.
- Real-time reporting means p0 arrives while it is still actionable. Watching the one-off rate during fieldwork tells you whether to keep recruiting; receiving it after analysis tells you what you should have done. Legacy survey platforms such as SurveyMonkey structure the work so that the answer can only arrive too late.
- Customizable AI consultants keep merge discipline. A consultant briefed to preserve and flag rare mentions rather than compress them is a policy you can apply to every study, instead of a rule you hope every analyst remembers on a Friday afternoon.
Common mistakes
- Computing f1 after the merge. The number will be near zero and it will mean nothing. Compute pre-merge, always.
- Counting interviews instead of mentions. p0 = f1/n uses coded mentions. Using participants inflates the estimate substantially.
- Reading Chao1 as the answer. It is a lower bound on the theme space, and a widely used one precisely because it is conservative.
- Promoting every singleton to a finding. The estimate does not say each one-off is important. It says the population contains kinds you have not met, which is a recruiting instruction, not a roadmap instruction.
- Reporting theme coverage without the mention coverage next to it. The pair is the insight; either number alone gets misread.
- Assuming a big corpus fixes it. Volume raises mention coverage quickly and theme coverage slowly. That gap is the reason a large study can be more confident and no better informed.
Frequently asked questions
What counts as a "mention" for the denominator?
One coded instance of one theme by one participant. If a participant raises the same theme four times in an interview, most teams count that once per participant per theme, which is the more conservative choice and keeps n from being inflated by talkative respondents. Whichever convention you pick, apply it before computing f1 and state it in the footer.
My f2 is zero. Does Chao1 break?
It switches form rather than breaking. When f2 = 0 the estimator becomes S-obs + f1(f1 - 1)/2, as given by Gotelli and Chao. An f2 of zero with a healthy f1 is itself a strong signal of a wide-open theme space: you are finding new kinds and not yet finding them twice.
Is a high singleton rate a sign of bad coding?
Sometimes, which is why the triage table matters. Splitter-style coding inflates f1 with artefacts. But a low f1 achieved by lumping is far more dangerous than a high f1 achieved by splitting, because lumping destroys the estimate while splitting only adds noise to it. Check agreement with inter-rater reliability before blaming the coder.
How does this relate to saturation?
It is the quantitative version of the same question. Data saturation asks whether new themes have stopped appearing; p0 estimates the rate at which they would still appear if you kept going, and 1* says whether the class is open. A study can look saturated and have 1* well below 1, which means the curve flattened for reasons other than coverage -- usually the recruiting channel.
Can I compute this on old studies?
Only if the mention-level data survived. This is the practical argument for a mention-level research repository rather than a folder of summary decks: reports preserve conclusions, and only raw coding preserves the ability to ask new questions of old data, including this one.
Does the estimate work for support tickets and reviews?
Yes, and often better, because volume is high and the coding is already mechanical. Treat each ticket's issue tag as a mention. Be careful with one thing: ticket systems have their own "Other" bucket and a queue that rewards closing tickets fast, so f1 is usually suppressed before you see it. Compute from raw text where you can.
Related Resources
- Capture-Recapture for Research -- the two-pass version of this measurement, when you can afford a second look.
- Why Your Theme Discovery Curve Flattens -- why an open theme space can still produce a flat curve.
- How to Code Qualitative Data -- where singletons are created and, usually, where they are lost.
- Affinity Mapping -- the synthesis step that merges hardest, and how to keep the pre-merge counts.
- Survivorship Bias in Customer Research -- the other way a corpus quietly stops representing the population.
- Structured Questions in AI Interviews -- the six question types that make the tail countable instead of invisible.
Related Articles
Affinity Mapping: Organize Qualitative Data Into Themes
Learn how to use affinity mapping to group qualitative research data into meaningful clusters and uncover actionable patterns.
How to Code Qualitative Data: A Step-by-Step Guide
Learn the complete process of qualitative coding — from building a codebook to identifying themes — and how AI tools like Koji automate the most time-consuming parts.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story
Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.
Activating Research Insights: Turn Findings Into Product Decisions
A practical guide to insight activation — the discipline of ensuring research findings actually drive product decisions. Covers why 40-60% of insights are never used, the 4-stage activation framework, decision-ready report formats, and how AI-native research platforms close the loop in real time.
How to Analyze Open-Ended Survey Responses with AI (2026 Guide)
Stop manually coding free-text survey responses. Learn how AI analyzes open-ended answers at scale — surfacing themes, sentiment, and quotes in minutes, plus why an AI interview captures 10x more depth than any survey can.