{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-22T15:54:50.094Z"},"content":[{"type":"documentation","id":"57a5badb-9330-4e72-814c-fde0fb4f3113","slug":"theme-discovery-curve-dependent-samples","title":"Why Your Theme Discovery Curve Flattens: Dependent Samples and the Illusion of Saturation (2026)","url":"https://www.koji.so/docs/theme-discovery-curve-dependent-samples","summary":"A flattening theme-discovery curve is jointly a property of the population and the sampling process, so it cannot distinguish coverage from a narrow recruiting channel. The US Census Bureau defines correlation bias as occurring when the probability of being included in one system influences the probability of inclusion in the other. In a worked case with 40 themes and two passes each reaching 60 percent, independence gives an accurate 84 percent coverage estimate while same-channel dependence reports 99 percent when true coverage is 66 percent. Van Rijnsoever (PLOS ONE, 2017) shows saturation depends more on the mean probability of observing codes than on how many codes exist.","content":"**Bottom line up front:** A theme-discovery curve that flattens is measuring two things at once -- how much of the population you have covered, and how narrow your recruiting is -- and it cannot tell you which one flattened it. When the two halves of a study come from the same channel, the same screener and the same week, the overlap between them is inflated by dependence, and every coverage estimate built on that overlap moves in the wrong direction. In a worked case where two passes each reach 60 percent of the theme space, independent sampling reports 84 percent coverage and is right; dependent sampling from a single channel reports **99 percent coverage while true coverage has fallen to 66 percent**. The curve got flatter and the study got worse. The fix is counterintuitive: to measure coverage, you have to make your second pass deliberately unlike your first.\n\n## The reflex this article exists to break\n\nOnce a team starts [estimating theme coverage from the overlap between two passes](/docs/capture-recapture-theme-coverage), the next instinct is a good one for almost every other purpose: make the passes comparable. Same screener, so the participants are alike. Same interview guide, so the prompts are alike. Same recruiting list, because that is the list that works. Same fortnight, so nothing in the product changes underneath.\n\nEvery one of those decisions increases the correlation between the two passes. And correlation between passes is the one thing that breaks the measurement, because the entire method rests on an assumption that the two looks are independent. Standardising your fieldwork is how you get a clean comparison and a corrupt estimate at the same time.\n\nThe US Census Bureau, which has run this exact problem on the largest scale anyone runs it, names the failure precisely. From the [Source and Accuracy of the 2020 Post-Enumeration Survey Estimates](https://www2.census.gov/programs-surveys/decennial/coverage-measurement/pes/2020-source-and-accuracy-pes-estimates.pdf): **\"Correlation bias occurs when the probability of being included in one system influences the probability of being included in the other system.\"** And it also arises, the same document notes, \"when there is heterogeneity in the capture probabilities of similar individuals, that is, when similar people have different probabilities of being included in either the census or the PES.\"\n\nThe people the census misses are disproportionately the people the follow-up survey misses too. Your research has the same structure. The customers who do not answer recruitment emails do not answer the second recruitment email either.\n\n## What the flattening curve is actually estimating\n\nA theme-discovery curve -- themes on the y-axis, interviews or coded mentions on the x-axis -- has a precise statistical meaning. [Gotelli and Chao](https://www.uvm.edu/~ngotelli/manuscriptpdfs/Gotelli_Chao_Encyclopedia_2013.pdf) state it exactly, writing about species accumulation curves in the *Encyclopedia of Biodiversity* (2013): \"the slope of the curve measured at any abundance level is **the probability that the next individual sampled represents a previously unsampled species**.\"\n\nThat is a genuinely useful quantity. Read it carefully, though, and the conditional is doing all the work. The slope estimates the probability that the next individual **sampled the way you have been sampling** is new. It is a property of the sampling process and the population jointly, and a flat curve is equally consistent with:\n\n- a population whose themes you have nearly enumerated, and\n- a sampling process that has stopped reaching new kinds of people.\n\nNothing inside the curve distinguishes them. This is why \"we hit saturation at 12\" is a statement about a recruiting pipeline at least as much as about a market.\n\n## The arithmetic of dependence\n\nHere is the mechanism with numbers attached. Suppose a population contains 40 distinct themes and each of two passes reaches 60 percent of them. Independence means a theme found in pass A has the same 60 percent chance of appearing in pass B as any other theme. Dependence means it has a higher chance -- because the same channel keeps returning the same kind of person.\n\n| Relationship between passes | P(found in B given found in A) | Overlap | Distinct themes observed | Estimated total | Reported coverage | True coverage |\n| --- | --- | --- | --- | --- | --- | --- |\n| Independent | 0.60 | 14.4 | 33.6 | 40.0 | 84.0 percent | 84.0 percent |\n| Mildly dependent | 0.75 | 18.0 | 30.0 | 32.0 | 93.8 percent | 75.0 percent |\n| Same channel, same screener | 0.90 | 21.6 | 26.4 | 26.7 | 99.0 percent | 66.0 percent |\n\nRead the last two columns together, because that pair is the whole point. As dependence rises, the number your study reports climbs from 84 to 99 percent while the coverage you actually have falls from 84 to 66 percent. The estimate does not get noisier as sampling gets narrower. **It gets confidently, monotonically wrong in the reassuring direction.** The bottom row is a study that finds fewer themes, sees more repetition, feels more saturated, and reports near-total coverage.\n\nThis is the same phenomenon the software-inspection literature ran into when it borrowed capture-recapture for defect counting. [Petersson, Thelin, Runeson and Wohlin](https://wohlin.eu/jss04-1.pdf), surveying ten years of that work in the *Journal of Systems and Software* (2004), list the assumption failures plainly, including that \"there is a risk that the reviewers co-operate, which violates the assumption of independence between reviewers,\" and conclude across the evaluation literature that **\"most estimators underestimate.\"**\n\n## Why the rare themes decide everything\n\nThe second half of the problem is that themes are not equally easy to observe, and the hard ones dominate the sample size you need. [Frank van Rijnsoever's simulation study](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0181689) in *PLOS ONE* (2017) quantifies this. Modelling a population of information sources holding codes, and defining theoretical saturation as observing every code at least once, he finds that **\"theoretical saturation is more dependent on the mean probability of observing codes than on the number of codes in a population.\"**\n\nThe magnitudes are startling. With 101 codes in the population and a mean probability of observing a code below 0.1, random-chance sampling \"generally requires more than **1000 sampling steps** to reach theoretical saturation in 95 percent of the cases.\" Under the same conditions, purposive sampling designed to yield at least one new code per step needs **about 46**, and maximally efficient purposive sampling about **20**.\n\nTwo conclusions follow, and they point the same way:\n\n1. **The tail is where the sample size lives.** Whether you need 20 informants or 1000 is decided by how hard your rarest codes are to observe, not by how many codes exist.\n2. **Purposive, deliberately varied sampling is not a compromise on rigour. It is between 20 and 50 times more efficient** than sampling at random from a single pool -- and it is also what restores the independence the coverage estimate needs.\n\nThe efficiency argument and the validity argument agree. That is rare enough to act on.\n\n## Diagnostics: is my curve flat for the right reason?\n\nFour checks, none of which need new fieldwork.\n\n**1. The channel restart test.** Plot the discovery curve separately for each recruiting channel, with the x-axis restarting at zero for each. If a channel that entered the study late produces a fresh burst of new themes, the pooled curve was flat because the earlier channel was exhausted, not because the population was.\n\n**2. Cross-channel theme lists.** Take the themes found only in channel A and only in channel B. If both exclusive lists are long, your channels are sampling different sub-populations and no single-channel curve was ever going to flatten honestly.\n\n**3. Screener drift.** Recruiters tighten screeners under deadline pressure, because looser screeners mean more no-shows. Compare the accepted profile in week 1 with week 3. A narrowing screener manufactures a flattening curve mechanically. See [research screener questions](/docs/research-screener-questions) for the failure modes.\n\n**4. The non-participant check.** Everyone in your sample said yes to something. [Nonresponse bias](/docs/nonresponse-bias) and [survivorship bias](/docs/survivorship-bias-customer-research) both act on the theme space, not just on the numbers, and they act on the second pass exactly as they acted on the first -- which is correlation bias in its purest form.\n\n## What to do instead\n\nDesign the second pass to be as unlike the first as the research question permits, then measure overlap across that gap.\n\n| Dimension | Pass 1 | Pass 2 should differ by |\n| --- | --- | --- |\n| Recruiting channel | In-product intercept | Community, panel, or customer-success referral |\n| Population slice | Active users | Churned users, evaluators who did not buy, blocked non-users |\n| Timing | Launch week | A quiet week, or after a support spike |\n| Moderator or brief | Standard discovery guide | A differently briefed consultant with different priors |\n| Modality | Text | Voice, which surfaces what people will not type |\n\nTwo practical warnings. First, if you vary the population slice, you have changed the population, and the closed-population assumption goes with it -- report coverage per slice rather than pooling. Second, do not vary the *matching rule*: how you decide two codes are the same theme must stay fixed, or the overlap becomes an artefact of coding style rather than of sampling. That distinction is the one thing to standardise. See [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research).\n\n## How Koji helps\n\nEvery prescription above costs calendar time in a traditional workflow, which is why teams standardise instead: one channel, one guide, one moderator, because that is what fits in three weeks. AI-native research removes the constraint that makes the wrong choice rational.\n\n- **Parallel fieldwork across channels.** Koji's AI-moderated interviews run concurrently rather than one scheduled slot at a time, so an in-product pass and a community pass can run in the same week instead of consecutive months. Independence stops being a luxury when it costs no extra calendar.\n- **Customizable AI consultants.** Different consultant briefs give genuinely different probing behaviour against the same population -- the closest practical analogue to a second, independent interviewer, without the second interviewer's schedule.\n- **Voice and text modalities.** Running part of a study by voice attacks the equal-catchability problem directly rather than acknowledging it in a limitations section.\n- **Automatic thematic analysis with per-source attribution.** Because Koji tags every theme with the interview and channel it came from, per-channel discovery curves and cross-channel exclusive theme lists are queries, not weekend projects. That is diagnostic 1 and diagnostic 2 done at no cost.\n- **Structured questions hold the instrument constant while the sample varies.** The six types -- open_ended, scale, single_choice, multiple_choice, ranking, and yes_no -- keep the closed part of the instrument identical across channels so that differences in the open-ended themes are attributable to the sample rather than to the questionnaire ([structured questions guide](/docs/structured-questions-guide)).\n- **Real-time reporting.** A curve that is still moving is actionable; a curve delivered with the readout is a post-mortem. Traditional survey tools such as SurveyMonkey give you the data when fieldwork closes, which is after the last moment you could have opened a second channel.\n\n## Common mistakes\n\n- **Pooling channels into one curve.** The pooled curve always flattens earlier than any component curve deserves.\n- **Reusing the same panel for the confirmatory study.** The confirmation is then almost guaranteed, and it means nothing.\n- **Treating high overlap as a quality signal.** It is a signal about your recruiting, not your rigour.\n- **Recruiting the second pass from referrals of the first.** Snowball sampling maximises dependence by construction; it is a fine discovery tactic and a terrible basis for a coverage estimate.\n- **Stopping at the flat spot without recording why.** Write down which channel and which screener produced the flatness. Six months later that note is the difference between a defensible claim and a folk memory.\n\n## Frequently asked questions\n\n### Does this mean saturation is a useless concept?\n\nNo. Saturation is a reasonable stopping heuristic and a poor coverage claim. Use it to decide when to stop spending on a channel; do not use it to assert that the theme space is enumerated. The two statements sound alike and differ by everything.\n\n### How different is different enough for the second pass?\n\nA useful test: could a participant plausibly have been recruited into either pass? If yes, the passes overlap by construction and dependence is high. Channels that share a source list -- your CRM, one community, one panel provider -- are the same channel wearing two names.\n\n### What if I only have one recruiting channel?\n\nThen report coverage as conditional on that channel, explicitly: \"within customers reachable by in-product prompt, estimated theme coverage was 90 percent.\" That sentence is honest, useful, and still tells a reader what the study cannot speak to. Splitting a single channel in half does not buy independence, though splitting by time can help a little if the population turns over.\n\n### Can a discovery curve rise again after flattening?\n\nYes, and it is the single most useful diagnostic there is. A curve that jumps when a new channel opens proves the earlier flat section was about the channel. This is why per-channel restarts belong in every report where a saturation claim is made.\n\n### Is correlation bias the same as sampling bias?\n\nThey are related but not identical. [Sampling bias](/docs/sampling-bias-research) makes your sample unrepresentative of the population. Correlation bias makes your *estimate of how unrepresentative you are* too optimistic, because the second look inherits the first look's blind spots. You can have unbiased-looking demographics and severe correlation bias at the same time.\n\n### How many channels are enough?\n\nTwo genuinely independent ones are enough to compute an estimate; three let you detect that two of them were secretly the same. If a third channel produces a long list of exclusive themes, treat the earlier two-channel estimate as a floor and say so.\n\n## Related Resources\n\n- [Capture-Recapture for Research](/docs/capture-recapture-theme-coverage) -- the estimator this article stress-tests, and how to compute it.\n- [The One-Off Comments Are the Only Estimate You Have](/docs/singleton-themes-unseen-coverage) -- what to do when you have only one pass and cannot measure overlap at all.\n- [How Many User Interviews Do You Need?](/docs/how-many-user-interviews) -- the sample-size guidance this refines.\n- [Sampling Bias: Types, Examples, and How to Avoid It](/docs/sampling-bias-research) -- the parent failure this article gives a measurement to.\n- [Screener Questions for User Research](/docs/screener-questions-guide) -- where curve-flattening quietly begins.\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) -- holding the instrument fixed while the sample varies.\n","category":"Research Methods","lastModified":"2026-08-22T03:24:38.720257+00:00","metaTitle":"Why Theme Discovery Curves Flatten: Dependent Samples (2026)","metaDescription":"Standardised recruiting inflates overlap and makes coverage estimates confidently wrong. Correlation bias, the channel restart test, and how to fix it.","keywords":["theme discovery curve","saturation curve flattening","dependent samples qualitative research","correlation bias","recruiting channel bias","qualitative coverage estimate","why no new themes"],"aiSummary":"A flattening theme-discovery curve is jointly a property of the population and the sampling process, so it cannot distinguish coverage from a narrow recruiting channel. The US Census Bureau defines correlation bias as occurring when the probability of being included in one system influences the probability of inclusion in the other. In a worked case with 40 themes and two passes each reaching 60 percent, independence gives an accurate 84 percent coverage estimate while same-channel dependence reports 99 percent when true coverage is 66 percent. Van Rijnsoever (PLOS ONE, 2017) shows saturation depends more on the mean probability of observing codes than on how many codes exist.","aiPrerequisites":["Familiarity with theme discovery or saturation curves","An understanding of how participants were recruited"],"aiLearningOutcomes":["Explain why a flat curve cannot prove coverage","Quantify how dependence biases a coverage estimate","Run the channel restart test and three other diagnostics","Design a second pass that restores independence"],"aiDifficulty":"advanced","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}