{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-06T14:10:25.693Z"},"content":[{"type":"documentation","id":"48f919b3-b0b7-41d0-8580-a65fc4429585","slug":"theme-dispersion-vs-mention-frequency","title":"Theme Dispersion: Why 47 Mentions Can Mean Three Customers (2026)","url":"https://www.koji.so/docs/theme-dispersion-vs-mention-frequency","summary":"Mention counts collapse volume and dispersion into one number. A theme with 47 mentions may come from 41 participants or from 3, and those corpora demand opposite decisions. Report distinct speakers and the largest contributor share beside every count; compute Gries DP (sum of |v_i - s_i| divided by 2) when a decision is expensive. Keep dispersion and frequency in separate columns, because Gries 2021 shows most dispersion measures blend the two. Dispersion depends strongly on the unit measured across, so choose per participant, per account or per segment to match the decision.","content":"**Bottom line up front:** A mention count is a numerator with a hidden internal structure. \"Forty-seven mentions of slow exports\" can mean forty-one different customers each raising it once, or three power users raising it sixteen times each. Those two corpora demand opposite decisions, and the count cannot tell them apart. Report the number of distinct speakers alongside every mention count, and add a dispersion statistic when the decision is expensive. Corpus linguistics has measured this for decades under the name dispersion, and the measures are well understood even though they are, in Gries's own assessment, neither widely known nor applied.\n\n## The quantity you are actually reporting\n\nWhen your analysis says a theme has 47 mentions, it has summed coded instances. That sum collapses two independent facts into one number:\n\n- **Volume** - how many times the theme was expressed.\n- **Dispersion** - how evenly those expressions were spread across the people who could have expressed them.\n\nVolume alone is the number almost every feedback tool surfaces by default, because it is the easy aggregate. It is also the one that drives roadmap arguments, which is precisely why its ambiguity is costly. A theme concentrated in a handful of unusually talkative participants looks, in a bar chart, exactly like a theme that everybody raised.\n\nStefan Gries made this the central argument of a 2008 paper in the *International Journal of Corpus Linguistics*. His point was that frequencies of occurrence and co-occurrence are the most frequent statistics in corpus linguistics, and that such frequencies in isolation may sometimes be misleading because they do not take into consideration the degree of dispersion of the item being counted. He proposed a deliberately simple measure, DP, to sit beside raw frequency rather than replace it.\n\n## A worked example where the count is identical\n\nTwo themes from a 60-interview study. Both have 47 coded mentions.\n\n| | Theme A: slow exports | Theme B: custom SSO |\n|---|---|---|\n| Coded mentions | 47 | 47 |\n| Distinct participants | 41 of 60 | 3 of 60 |\n| Mentions per participant | 1.1 | 15.7 |\n| Largest single contributor | 3 mentions (6%) | 24 mentions (51%) |\n\nTheme A is a broad, shallow irritation: two thirds of your sample hit it, nobody dwells on it. Theme B is a narrow, deep obsession: three accounts, one of them supplying half the evidence on its own.\n\nBoth are real. Both may deserve work. But they are different findings with different economics, and \"47 mentions\" is the one summary that makes them indistinguishable. If Theme B's three participants happen to be your three largest accounts, it may well outrank Theme A. If they are three trial users who churned anyway, it should not. **The count cannot carry that argument; the dispersion can.**\n\n## How to compute dispersion without a linguistics degree\n\nStart with the two numbers you can produce today from any coded corpus:\n\n1. **Distinct speaker count.** How many different participants raised the theme at least once. This is the single highest-value column missing from most theme tables.\n2. **Concentration.** The share of the theme's mentions contributed by its single largest contributor. Above roughly 30 percent, treat the theme as one person's view until proven otherwise.\n\nWhen a decision justifies more rigour, compute Gries's DP. Split your corpus into parts (see the unit problem below). For each part *i*, let *s_i* be that part's share of the whole corpus and *v_i* be that part's share of the theme's mentions. Then:\n\n**DP = ( sum of | v_i - s_i | ) / 2**\n\nDP runs from 0, meaning the theme is spread exactly in proportion to where it could have appeared, up to close to 1, meaning it is confined to a vanishingly small slice. Lower is more even. The arithmetic is deliberately undemanding: it is a sum of absolute differences, halved.\n\nOne important caveat before you lean on any such number. In a 2021 paper in the *Journal of Second Language Studies*, Gries revisited the field's most widely used dispersion measures and argued that most of them are not particularly valid, in the sense that they measure an amalgam of a lot of frequency and a little dispersion rather than dispersion itself. The practical lesson is not to abandon dispersion. It is to keep dispersion and frequency in **separate columns** and refuse to blend them into a single composite score, because a blended score quietly reintroduces the ambiguity you were trying to remove.\n\n## The unit problem, which is where most teams go wrong\n\nDispersion is not a property of a theme alone. It is a property of a theme measured across a chosen set of units. Change the units and the answer changes.\n\nEgbert, Burch and Biber demonstrated this directly in a 2020 *International Journal of Corpus Linguistics* study applying a dispersion index built for unequal-sized corpus parts to the British National Corpus. Their finding was that the dispersion of a word is strongly influenced by the corpus units or parts it is measured across, and their recommendation was that dispersion should be measured and interpreted based on corpus units that are linguistically meaningful for the particular research goal.\n\nTranslate that into customer research and it becomes the most actionable idea in this article. Your candidate units are:\n\n- **Per interview.** The default, and usually the wrong one. If one account gave you nine interviews, a theme confined to that account looks well dispersed.\n- **Per participant.** Better. Removes the talkative-individual artefact.\n- **Per account or company.** Correct for most B2B decisions, because the buying and churn unit is the account, not the seat.\n- **Per segment.** Correct when you are deciding whether something is a universal problem or a segment problem.\n\nA theme can be well dispersed per interview and badly concentrated per account at the same time, and both statements are true. **Pick the unit that matches the decision you are about to make, state it, and keep it stable across reports** so that month-to-month comparisons mean something.\n\n## How this differs from the missing-denominator problem\n\nIt is worth separating two failures that look similar and are not.\n\nThe missing-denominator problem, covered in [why complaint counts cannot become rates](/docs/complaint-counts-cannot-be-rates), is about the **outside** of the fraction: you cannot convert a count into a rate without knowing how many people were exposed and could have complained. Dispersion is about the **inside** of the numerator: given the mentions you did collect, how concentrated are they among the people who produced them.\n\nThese are independent. A theme can have a perfectly good denominator and still be three people repeating themselves. A theme can be beautifully dispersed across your sample and still tell you nothing about prevalence in your user base, because your sample was self-selected. Fixing one does not fix the other, and a theme table that addresses neither is a table of numbers that cannot support a decision.\n\n## How Koji handles this\n\nKoji is built so that the speaker behind every mention is never lost, which is the precondition for measuring dispersion at all. Several design choices matter here:\n\n- **Stable question IDs.** Koji's study questions carry stable identifiers that preserve traceability from interview plan, through the AI interviewer, into analysis, and on to report aggregation. A theme in a Koji report remains attributable to the specific participants and the specific question that produced it, so a distinct-speaker count is always recoverable rather than something you reconstruct by hand from a spreadsheet.\n- **Structured questions fix the closed part of the instrument.** Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - as first-class definitions rather than free text. Because the closed questions are fixed and identically asked, your segment and account attributes arrive clean, which is what lets you switch the dispersion unit from interview to participant to account without re-coding anything. See the [structured questions guide](/docs/structured-questions-guide) for how the six types behave.\n- **Consistent probing across participants.** In a manual study, a theme's mention count partly measures which interviewer chose to dig. Koji's AI interviewer applies the same follow-up logic to every participant, so differences in mention volume reflect the participants rather than the moderator's energy on a Friday afternoon.\n- **Quality scores you can filter on.** Every Koji interview receives a 1-5 quality score with a breakdown across relevance, depth and coverage. Before you trust a concentration figure, filter out the thin interviews: a theme that looks concentrated may simply be the only theme that survived in a set of shallow conversations.\n- **Voice and text in one corpus.** Koji runs both voice and text interviews, and mention volume differs systematically by mode because people speak more words than they type. If your corpus mixes modes, compute dispersion per participant rather than per mention so the mode does not masquerade as enthusiasm.\n\nThe practical workflow: let Koji's automatic theme extraction propose the themes, then read every theme with its distinct-speaker count beside it. Koji's reports keep that link intact, so the question \"how many different people actually said this\" takes seconds rather than an afternoon of transcript archaeology.\n\n## Common mistakes\n\n- **Reporting mention counts with no speaker count.** The single most common and most expensive omission. Add one column.\n- **Deduplicating to one mention per person and discarding the rest.** This solves concentration by destroying intensity. Keep both numbers; a person who raises something nine times is telling you something real.\n- **Blending frequency and dispersion into one score.** Produces a number that is mostly frequency wearing a disguise, which is Gries's 2021 objection.\n- **Letting the unit drift between reports.** Per-interview one month and per-account the next makes every trend meaningless.\n- **Treating high concentration as automatic disqualification.** Concentrated themes are how you find the account that is about to churn. Concentration is a flag for interpretation, not a reason to delete.\n\n## Frequently asked questions\n\n### What is theme dispersion in customer feedback analysis?\n\nTheme dispersion measures how evenly a theme's mentions are spread across the participants, accounts or segments that could have raised it, as opposed to how many mentions it received in total. A theme with 47 mentions from 41 people is highly dispersed; a theme with 47 mentions from 3 people is highly concentrated. The concept comes from corpus linguistics, where Gries's 2008 DP measure is the best known simple index, and it answers a question raw frequency cannot: is this everybody's problem or a few people's problem?\n\n### How many mentions do I need before a theme is real?\n\nMention count is the wrong threshold to set, because it can be satisfied by one talkative participant. Set your threshold on distinct speakers instead. A practical default for product decisions is at least three distinct participants from at least two different accounts before a theme leaves the exploratory column, with the single-largest-contributor share under about 30 percent. A theme that fails those tests is not false, it is simply not yet evidence of a shared problem.\n\n### Should I count a theme once per person or once per mention?\n\nKeep both, in separate columns. Counting once per person removes the talkative-participant artefact and is the better basis for deciding whether a problem is widespread. Counting every mention preserves intensity, which is genuine information about how much the problem bothers people. Collapsing to one of the two throws away a dimension you will want later, and the pair together is what makes a theme table interpretable.\n\n### What is the difference between dispersion and the missing denominator problem?\n\nThey sit on opposite sides of the same fraction. The missing denominator is about not knowing how many people were exposed and could have mentioned something, which is why counts cannot become rates. Dispersion is about the internal structure of the mentions you did collect, specifically how concentrated they are among their speakers. A theme can fail one test and pass the other, so fixing a denominator does not tell you anything about concentration, and vice versa.\n\n### Which unit should I measure dispersion across?\n\nThe unit that matches the decision. For B2B roadmap and churn decisions, measure across accounts, because the account is what renews. For universality questions, measure across segments. Per participant is a reasonable general default; per interview is usually wrong, because one account contributing many interviews will make a narrow theme look broad. Egbert, Burch and Biber showed that a word's dispersion is strongly influenced by the units it is measured across, so the choice is not cosmetic, and whichever unit you pick should stay stable between reports.\n\n### Can I compute dispersion automatically in Koji?\n\nKoji preserves the link between every theme, the participant who raised it and the question that produced it, via stable question IDs that carry through from the interview plan to report aggregation. That is the hard part, and it is what makes distinct-speaker counts and concentration shares available directly from a Koji report rather than requiring manual transcript work. Combined with clean segment and account attributes from Koji's structured question types, you can switch the dispersion unit between participant, account and segment without re-coding the corpus.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types that keep your segment and account attributes clean enough to switch dispersion units\n- [Why Complaint Counts Cannot Become Rates](/docs/complaint-counts-cannot-be-rates) - the denominator problem that sits on the other side of the same fraction\n- [Singleton Themes and Unseen Coverage](/docs/singleton-themes-unseen-coverage) - what the themes mentioned exactly once tell you about the ones you never heard\n- [Why the Loudest Complaint Hides the Real One](/docs/dominant-complaint-masking-interviews) - how a concentrated theme suppresses the themes around it\n- [The Composite Sample Problem](/docs/composite-sample-aggregate-feedback-variance) - why an aggregate score hides the account that is on fire\n- [Feedback Volume Tracks Attention, Not Incidence](/docs/feedback-volume-reporting-propensity) - the time-series cousin of this problem\n","category":"Analysis & Synthesis","lastModified":"2026-10-06T07:51:42.491682+00:00","metaTitle":"Theme Dispersion vs Mention Frequency in Feedback (2026)","metaDescription":"A theme with 47 mentions can be 41 customers or 3. Add a distinct-speaker column, pick the right unit, and compute Gries DP when it matters.","keywords":["theme dispersion","mention count vs customer count","how many customers mentioned a theme","distinct speaker count feedback","gries dp dispersion","feedback theme frequency analysis"],"aiSummary":"Mention counts collapse volume and dispersion into one number. A theme with 47 mentions may come from 41 participants or from 3, and those corpora demand opposite decisions. Report distinct speakers and the largest contributor share beside every count; compute Gries DP (sum of |v_i - s_i| divided by 2) when a decision is expensive. Keep dispersion and frequency in separate columns, because Gries 2021 shows most dispersion measures blend the two. Dispersion depends strongly on the unit measured across, so choose per participant, per account or per segment to match the decision.","aiPrerequisites":["A corpus of coded themes with participant attribution","Basic familiarity with thematic analysis"],"aiLearningOutcomes":["Separate mention volume from theme dispersion","Compute distinct-speaker counts, concentration shares and Gries DP","Choose a dispersion unit that matches the decision at hand","Distinguish dispersion from the missing-denominator problem"],"aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}