Back to docs
Analysis & Synthesis

Theme Dispersion: Why 47 Mentions Can Mean Three Customers (2026)

A mention count hides how many people produced it. 47 mentions can be 41 customers or 3. Report distinct speakers beside every count, choose your dispersion unit deliberately, and use Gries DP when the decision is expensive.

Bottom line up front: A mention count is a numerator with a hidden internal structure. "Forty-seven mentions of slow exports" can mean forty-one different customers each raising it once, or three power users raising it sixteen times each. Those two corpora demand opposite decisions, and the count cannot tell them apart. Report the number of distinct speakers alongside every mention count, and add a dispersion statistic when the decision is expensive. Corpus linguistics has measured this for decades under the name dispersion, and the measures are well understood even though they are, in Gries's own assessment, neither widely known nor applied.

The quantity you are actually reporting

When your analysis says a theme has 47 mentions, it has summed coded instances. That sum collapses two independent facts into one number:

  • Volume - how many times the theme was expressed.
  • Dispersion - how evenly those expressions were spread across the people who could have expressed them.

Volume alone is the number almost every feedback tool surfaces by default, because it is the easy aggregate. It is also the one that drives roadmap arguments, which is precisely why its ambiguity is costly. A theme concentrated in a handful of unusually talkative participants looks, in a bar chart, exactly like a theme that everybody raised.

Stefan Gries made this the central argument of a 2008 paper in the International Journal of Corpus Linguistics. His point was that frequencies of occurrence and co-occurrence are the most frequent statistics in corpus linguistics, and that such frequencies in isolation may sometimes be misleading because they do not take into consideration the degree of dispersion of the item being counted. He proposed a deliberately simple measure, DP, to sit beside raw frequency rather than replace it.

A worked example where the count is identical

Two themes from a 60-interview study. Both have 47 coded mentions.

Theme A: slow exportsTheme B: custom SSO
Coded mentions4747
Distinct participants41 of 603 of 60
Mentions per participant1.115.7
Largest single contributor3 mentions (6%)24 mentions (51%)

Theme A is a broad, shallow irritation: two thirds of your sample hit it, nobody dwells on it. Theme B is a narrow, deep obsession: three accounts, one of them supplying half the evidence on its own.

Both are real. Both may deserve work. But they are different findings with different economics, and "47 mentions" is the one summary that makes them indistinguishable. If Theme B's three participants happen to be your three largest accounts, it may well outrank Theme A. If they are three trial users who churned anyway, it should not. The count cannot carry that argument; the dispersion can.

How to compute dispersion without a linguistics degree

Start with the two numbers you can produce today from any coded corpus:

  1. Distinct speaker count. How many different participants raised the theme at least once. This is the single highest-value column missing from most theme tables.
  2. Concentration. The share of the theme's mentions contributed by its single largest contributor. Above roughly 30 percent, treat the theme as one person's view until proven otherwise.

When a decision justifies more rigour, compute Gries's DP. Split your corpus into parts (see the unit problem below). For each part i, let s_i be that part's share of the whole corpus and v_i be that part's share of the theme's mentions. Then:

DP = ( sum of | v_i - s_i | ) / 2

DP runs from 0, meaning the theme is spread exactly in proportion to where it could have appeared, up to close to 1, meaning it is confined to a vanishingly small slice. Lower is more even. The arithmetic is deliberately undemanding: it is a sum of absolute differences, halved.

One important caveat before you lean on any such number. In a 2021 paper in the Journal of Second Language Studies, Gries revisited the field's most widely used dispersion measures and argued that most of them are not particularly valid, in the sense that they measure an amalgam of a lot of frequency and a little dispersion rather than dispersion itself. The practical lesson is not to abandon dispersion. It is to keep dispersion and frequency in separate columns and refuse to blend them into a single composite score, because a blended score quietly reintroduces the ambiguity you were trying to remove.

The unit problem, which is where most teams go wrong

Dispersion is not a property of a theme alone. It is a property of a theme measured across a chosen set of units. Change the units and the answer changes.

Egbert, Burch and Biber demonstrated this directly in a 2020 International Journal of Corpus Linguistics study applying a dispersion index built for unequal-sized corpus parts to the British National Corpus. Their finding was that the dispersion of a word is strongly influenced by the corpus units or parts it is measured across, and their recommendation was that dispersion should be measured and interpreted based on corpus units that are linguistically meaningful for the particular research goal.

Translate that into customer research and it becomes the most actionable idea in this article. Your candidate units are:

  • Per interview. The default, and usually the wrong one. If one account gave you nine interviews, a theme confined to that account looks well dispersed.
  • Per participant. Better. Removes the talkative-individual artefact.
  • Per account or company. Correct for most B2B decisions, because the buying and churn unit is the account, not the seat.
  • Per segment. Correct when you are deciding whether something is a universal problem or a segment problem.

A theme can be well dispersed per interview and badly concentrated per account at the same time, and both statements are true. Pick the unit that matches the decision you are about to make, state it, and keep it stable across reports so that month-to-month comparisons mean something.

How this differs from the missing-denominator problem

It is worth separating two failures that look similar and are not.

The missing-denominator problem, covered in why complaint counts cannot become rates, is about the outside of the fraction: you cannot convert a count into a rate without knowing how many people were exposed and could have complained. Dispersion is about the inside of the numerator: given the mentions you did collect, how concentrated are they among the people who produced them.

These are independent. A theme can have a perfectly good denominator and still be three people repeating themselves. A theme can be beautifully dispersed across your sample and still tell you nothing about prevalence in your user base, because your sample was self-selected. Fixing one does not fix the other, and a theme table that addresses neither is a table of numbers that cannot support a decision.

How Koji handles this

Koji is built so that the speaker behind every mention is never lost, which is the precondition for measuring dispersion at all. Several design choices matter here:

  • Stable question IDs. Koji's study questions carry stable identifiers that preserve traceability from interview plan, through the AI interviewer, into analysis, and on to report aggregation. A theme in a Koji report remains attributable to the specific participants and the specific question that produced it, so a distinct-speaker count is always recoverable rather than something you reconstruct by hand from a spreadsheet.
  • Structured questions fix the closed part of the instrument. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - as first-class definitions rather than free text. Because the closed questions are fixed and identically asked, your segment and account attributes arrive clean, which is what lets you switch the dispersion unit from interview to participant to account without re-coding anything. See the structured questions guide for how the six types behave.
  • Consistent probing across participants. In a manual study, a theme's mention count partly measures which interviewer chose to dig. Koji's AI interviewer applies the same follow-up logic to every participant, so differences in mention volume reflect the participants rather than the moderator's energy on a Friday afternoon.
  • Quality scores you can filter on. Every Koji interview receives a 1-5 quality score with a breakdown across relevance, depth and coverage. Before you trust a concentration figure, filter out the thin interviews: a theme that looks concentrated may simply be the only theme that survived in a set of shallow conversations.
  • Voice and text in one corpus. Koji runs both voice and text interviews, and mention volume differs systematically by mode because people speak more words than they type. If your corpus mixes modes, compute dispersion per participant rather than per mention so the mode does not masquerade as enthusiasm.

The practical workflow: let Koji's automatic theme extraction propose the themes, then read every theme with its distinct-speaker count beside it. Koji's reports keep that link intact, so the question "how many different people actually said this" takes seconds rather than an afternoon of transcript archaeology.

Common mistakes

  • Reporting mention counts with no speaker count. The single most common and most expensive omission. Add one column.
  • Deduplicating to one mention per person and discarding the rest. This solves concentration by destroying intensity. Keep both numbers; a person who raises something nine times is telling you something real.
  • Blending frequency and dispersion into one score. Produces a number that is mostly frequency wearing a disguise, which is Gries's 2021 objection.
  • Letting the unit drift between reports. Per-interview one month and per-account the next makes every trend meaningless.
  • Treating high concentration as automatic disqualification. Concentrated themes are how you find the account that is about to churn. Concentration is a flag for interpretation, not a reason to delete.

Frequently asked questions

What is theme dispersion in customer feedback analysis?

Theme dispersion measures how evenly a theme's mentions are spread across the participants, accounts or segments that could have raised it, as opposed to how many mentions it received in total. A theme with 47 mentions from 41 people is highly dispersed; a theme with 47 mentions from 3 people is highly concentrated. The concept comes from corpus linguistics, where Gries's 2008 DP measure is the best known simple index, and it answers a question raw frequency cannot: is this everybody's problem or a few people's problem?

How many mentions do I need before a theme is real?

Mention count is the wrong threshold to set, because it can be satisfied by one talkative participant. Set your threshold on distinct speakers instead. A practical default for product decisions is at least three distinct participants from at least two different accounts before a theme leaves the exploratory column, with the single-largest-contributor share under about 30 percent. A theme that fails those tests is not false, it is simply not yet evidence of a shared problem.

Should I count a theme once per person or once per mention?

Keep both, in separate columns. Counting once per person removes the talkative-participant artefact and is the better basis for deciding whether a problem is widespread. Counting every mention preserves intensity, which is genuine information about how much the problem bothers people. Collapsing to one of the two throws away a dimension you will want later, and the pair together is what makes a theme table interpretable.

What is the difference between dispersion and the missing denominator problem?

They sit on opposite sides of the same fraction. The missing denominator is about not knowing how many people were exposed and could have mentioned something, which is why counts cannot become rates. Dispersion is about the internal structure of the mentions you did collect, specifically how concentrated they are among their speakers. A theme can fail one test and pass the other, so fixing a denominator does not tell you anything about concentration, and vice versa.

Which unit should I measure dispersion across?

The unit that matches the decision. For B2B roadmap and churn decisions, measure across accounts, because the account is what renews. For universality questions, measure across segments. Per participant is a reasonable general default; per interview is usually wrong, because one account contributing many interviews will make a narrow theme look broad. Egbert, Burch and Biber showed that a word's dispersion is strongly influenced by the units it is measured across, so the choice is not cosmetic, and whichever unit you pick should stay stable between reports.

Can I compute dispersion automatically in Koji?

Koji preserves the link between every theme, the participant who raised it and the question that produced it, via stable question IDs that carry through from the interview plan to report aggregation. That is the hard part, and it is what makes distinct-speaker counts and concentration shares available directly from a Koji report rather than requiring manual transcript work. Combined with clean segment and account attributes from Koji's structured question types, you can switch the dispersion unit between participant, account and segment without re-coding the corpus.

Related Resources

Related Articles

Capture-Recapture for Research: How to Estimate the Themes Your Study Never Found (2026)

Two independent coding passes turn "no new themes" into a number: the overlap between them estimates how many themes neither pass ever reached.

How to Code Qualitative Data: A Step-by-Step Guide

Learn the complete process of qualitative coding — from building a codebook to identifying themes — and how AI tools like Koji automate the most time-consuming parts.

Why Complaint Counts Cannot Become Rates (And What to Compute Instead)

A count of complaints has no denominator, so it can never become a rate. Here is the arithmetic that works anyway, borrowed from fifty years of safety surveillance.

The Composite Sample Problem: Why Your Aggregate Score Hides the Account on Fire (2026)

Combining many responses into one number gives you an unbiased mean and destroys the between-unit variance. Here is what compositing costs you, and the design that keeps the mean and the outlier.

Customer Feedback Categorization: How to Build a Feedback Taxonomy That Scales (2026)

A practical guide to categorizing customer feedback: designing a feedback taxonomy, choosing flat vs. hierarchical tags, avoiding tag sprawl, and using AI auto-tagging to turn thousands of unstructured comments into quantified themes.

Why the Loudest Complaint Hides the Real One: Masking in Customer Interviews (2026)

One dominant complaint does not just take up airtime - it raises the threshold for everything quieter, asymmetrically, and your analysis then discards what was buried. A protocol for hearing the masked signal.

Feedback Volume Tracks Attention, Not Incidence: How to Read a Complaint Trend

A rise or fall in complaint volume is at least as likely to be a change in how willing people are to report as a change in your product. Here is how to tell them apart.

Singleton Themes: Why One-Off Comments Are the Only Estimate You Have of What You Missed (2026)

Good-Turing says the chance the next respondent raises something new is the singleton count divided by total mentions. Every synthesis step deletes singletons first.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

The Complete Guide to Thematic Analysis

Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.