Back to docs
Analysis & Synthesis

Keyness: Which Words Are Actually Distinctive to a Customer Segment (2026)

Word frequency tells you what is common; keyness tells you what is distinctive. Choose a disjoint reference corpus, count each term once per document to avoid the dispersion trap Egbert and Biber documented, and read the passages before interpreting.

Bottom line up front: If you compare what churned customers said against what retained customers said by counting words, you will rediscover that both groups say "the", "product" and "team" a lot. The statistic you actually want is keyness: how much more characteristic a term is of one group than of a reference group. And there is a well-documented trap inside it. Standard keyness is computed over a corpus treated as one undifferentiated block, so it rewards terms that a few long, verbose transcripts happened to repeat. Compute keyness on the number of documents containing a term, not the number of occurrences, and the list stops being an artefact of your three chattiest participants.

What keyness is, in one paragraph

Keyness compares two corpora: a target corpus (the group you care about, say the 40 interviews with customers who churned) and a reference corpus (the comparison group, say the 180 interviews with customers who renewed). For each term, you test whether its relative frequency in the target is higher than you would expect given its relative frequency in the reference. The conventional test statistic is log-likelihood, popularised for this family of problems by Ted Dunning's 1993 work on the statistics of surprise and coincidence in Computational Linguistics; mutual information, introduced to lexicography by Church and Hanks at the 27th Annual Meeting of the Association for Computational Linguistics, is the other long-standing option.

The output is a ranked list of terms that are over-represented in the target. That list is the closest thing text analysis has to an answer to "what is different about these people".

The reference corpus decides the answer

This is the part that catches everyone, and it has no statistical fix. Keyness is not a property of your target corpus. It is a property of the pair. Change the reference and the keywords change.

Run the same 40 churned-customer interviews against three different references:

Reference corpusWhat rises to the topWhat it tells you
Retained customerspricing, approval, budgetWhy these customers left rather than stayed
All customers, churned includedmild versions of the same terms, weaker signalDiluted, because the target is inside the reference
General Englishproduct, dashboard, onboardingThat you work in software. Useless

The third row is the common own-goal: comparing customer interviews against a general-language baseline surfaces your own product vocabulary, which you already knew. The second row is the subtler one. If your target is a subset of your reference, you are comparing a group against itself plus others, which shrinks every difference toward zero. Keep the two corpora disjoint.

State your reference corpus in every keyness result you circulate. A keyword list without its reference is uninterpretable, in the same way that a percentage without a denominator is uninterpretable.

The dispersion trap inside keyness

Here is the failure that the corpus linguistics literature has documented and that almost no feedback tool accounts for.

Keyword analysis has become a standard tool for identifying the words especially characteristic of texts in a target domain. But as Jesse Egbert and Douglas Biber set out in a 2019 study in Corpora, the statistical computation of keyness makes no reference to those texts. Once the corpus has been constructed, it is treated as a homogeneous whole for the purpose of computing keyness. The consequence they identify is precise: the keywords in such lists are relatively frequent in the corpus, but they are often not widely dispersed across the texts of that corpus, and are therefore not truly representative of the target discourse domain.

In feedback terms: one customer who mentions "latency" thirty times across a long interview can push "latency" to the top of your churn keyword list, even though no other churned customer raised it at all. The statistic cannot see that the thirty occurrences came from one person, because the corpus was flattened before the test ran.

Egbert and Biber's proposed remedy is text dispersion keyness, which computes keyness from the number of texts in which a term appears rather than from its corpus frequency. They compared it against four other methods for computing keyness in a series of case studies identifying the keywords typical of online travel blogs, assessing each method on content-generalisability and content-distinctiveness, and concluded that text dispersion keyness is the superior measure for generating keyword lists.

The practitioner version is almost embarrassingly simple:

Count each term at most once per document. Then run your keyness test on those document counts.

A term now scores highly only if many different interviews contain it. Repetition within one interview stops counting. You can implement this with a single change to how you build the input table, and it removes an entire class of false keyword.

A worked comparison

Forty churned interviews against 180 retained interviews. The same corpus, scored two ways.

TermOccurrences (churn)Interviews containing itRank by occurrenceRank by document count
pricing6128 of 4021
latency343 of 40319
approval4422 of 4042
migration775 of 40114

Ranked by occurrence, your churn story is about migration and latency. Ranked by document count, it is about pricing and approval. The second story is the one that generalises: migration dominated the raw count because two customers described a migration at length, and a theme present in 5 of 40 interviews is not what characterises the group.

Both columns are worth keeping. A high-occurrence, low-document term like migration is a genuine signal about a small number of accounts, and it belongs in your notes. It simply does not belong at the top of a list labelled "what churned customers talk about".

Choosing and reading a keyness list

A workable procedure:

  1. Define the two corpora and keep them disjoint. Target and reference must not overlap.
  2. Normalise the comparison for size. The groups will differ in interview count and length; the test statistic handles this, but only if you feed it correct totals.
  3. Count once per document. Per the dispersion fix above.
  4. Strip your own product vocabulary. Your brand name, feature names and the words from your own question wording will dominate otherwise. Participants echo the question they were asked.
  5. Read the top 30 in context before interpreting any of them. A keyword is a pointer to passages, not a finding. "Approval" could mean procurement friction or a praised approvals feature; only the surrounding text distinguishes them.
  6. Report the reference corpus alongside the list. Always.

Step 4 deserves emphasis because it interacts with how you collect data. If your interview guide asks about pricing, "pricing" will be key in every corpus you build, and you will have discovered your own script.

How Koji handles this

Keyness analysis needs two things that are hard to get from a pile of transcripts: clean group membership, and the ability to count per document rather than per occurrence. Koji provides both.

  • Clean group labels come from structured questions. Koji supports six first-class structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - so plan, segment, tenure and renewal intent arrive as structured values rather than something you infer from prose. That is what makes a disjoint target and reference corpus definable in the first place. The structured questions guide covers how each type is asked and aggregated.
  • Stable question IDs keep the document boundary intact. Every Koji study question carries a stable identifier that preserves traceability from the interview plan through the AI interviewer to analysis and report aggregation. Because each response stays attached to its participant and its question, counting a term once per interview, or once per account, is a grouping choice rather than a data-recovery project.
  • Comparable instruments across groups. A keyness comparison is only valid if both groups were asked comparable questions. Koji's AI interviewer works from the same study definition for every participant, so your churned and retained corpora differ because the customers differ, not because a human moderator asked the churn cohort a pointed extra question.
  • Insights chat for the context step. Step 5 above, reading keywords in context, is the step teams skip. Koji lets you query the corpus conversationally and pull the passages behind a term, so checking whether "approval" means friction or praise takes one question instead of a transcript search.
  • Quality scores to screen the input. Koji scores each interview 1-5 across relevance, depth and coverage. Thin interviews contribute few terms and distort document counts; filtering them before a keyness run is a two-second change in Koji and a meaningful improvement in the output.

Used this way, Koji turns a keyness analysis from a scripting exercise into a filter-and-read workflow, which matters because the value of keyness is almost entirely in the reading.

Common mistakes

  • Comparing against general English. Surfaces your own product vocabulary. Compare customers against customers.
  • Letting the target sit inside the reference. Dilutes every difference. Keep them disjoint.
  • Ranking by raw occurrence. Lets one verbose participant write your headline, which is the Egbert and Biber objection.
  • Treating a keyword as a finding. It is a pointer to passages. Read them.
  • Forgetting your question wording is in the corpus. You will find the words you asked about.
  • Circulating a list without its reference corpus. Nobody downstream can interpret it.

Frequently asked questions

What is keyness analysis?

Keyness analysis identifies the terms that are over-represented in one corpus relative to a reference corpus, using a statistical test such as log-likelihood. In customer research it answers questions like "what do churned customers talk about that retained customers do not". The output is a ranked list of distinctive terms rather than common ones, which is what distinguishes it from a word frequency count or a word cloud.

How is keyness different from just counting word frequency?

Word frequency tells you what is common; keyness tells you what is distinctive. Frequency lists from any two customer groups look nearly identical, because both are dominated by ordinary English and your own product vocabulary. Keyness divides out that shared baseline by comparing against a reference corpus, so what remains is the difference between the groups rather than the content they share.

Which reference corpus should I use?

Use the most similar group that excludes your target. To characterise churned customers, compare them against retained customers, not against all customers and not against general English. Comparing against general English returns your product vocabulary, and comparing against a superset that contains your target shrinks every difference toward zero. Whatever you choose, report it with the results, because a keyword list cannot be interpreted without knowing what it was compared against.

Why do my keyword lists keep surfacing things only one customer said?

Because standard keyness is computed over the corpus as a single undifferentiated block, so thirty repetitions by one talkative participant count the same as thirty mentions spread across thirty people. Egbert and Biber documented exactly this, noting that conventional keywords are frequent in the corpus but often not widely dispersed across its texts. The fix is to count each term at most once per interview and run the test on those document counts.

What is text dispersion keyness?

Text dispersion keyness computes keyness from the number of texts in which a term appears rather than from its total frequency in the corpus. Egbert and Biber proposed it in 2019, compared it against four other keyness methods across case studies on online travel blogs, and found it superior for generating keyword lists when judged on content-generalisability and content-distinctiveness. In practice it means counting once per document, which makes a term score highly only when many different documents contain it.

Can I run keyness analysis on Koji interview data?

Yes, and the two prerequisites are what Koji is built to supply. Group membership comes from Koji's structured question types, so you can define disjoint target and reference corpora from structured values rather than inferring them from prose. Stable question IDs keep every response attached to its participant and question, so you can count terms once per interview or once per account as needed. Koji's insights chat then covers the step that matters most, which is reading the passages behind each keyword before you interpret it.

Related Resources

Related Articles

Why a Tagged Quote Is Not Evidence: The Archival Bond in Research Repositories

A quote filed under a theme is information; the interview it came from is evidence. The five bonds repositories delete at intake, and how to keep them.

Content Analysis: The Complete Guide to Analyzing Text and Interview Data

A comprehensive guide to content analysis as a research method — covering conventional, directed, and summative approaches, step-by-step coding, inter-rater reliability, and how AI automates the most time-consuming parts.

Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)

Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.

NPS Comment Analysis: Turn Open-Text Verbatims into Driver Themes

How to analyze NPS open-text comments at scale — code verbatims into driver themes, link them to score movement, and use AI follow-up to capture the why behind every rating.

Has Your Open-End Data Gone Flat? Measuring Lexical Diversity Instead of Accusing Participants

You cannot judge whether one response was AI-written, but you can measure whether a whole batch has lost its variance. A practical guide to lexical diversity and near-duplicate monitoring for open-ended research data.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Theme Dispersion: Why 47 Mentions Can Mean Three Customers (2026)

A mention count hides how many people produced it. 47 mentions can be 41 customers or 3. Report distinct speakers beside every count, choose your dispersion unit deliberately, and use Gries DP when the decision is expensive.

Topic Modeling for Customer Feedback: How to Find Themes in Open-Ended Responses at Scale

A practical guide to topic modeling for customer feedback — how LDA and modern NLP surface hidden themes in open-ended survey responses and reviews, the limitations of traditional methods, and the faster AI-native alternative.

Verbatim Analysis: How to Code and Analyze Open-Ended Responses at Scale (2026)

Verbatim coding turns messy open-ended answers into countable themes — but manual coding is slow, expensive, and bias-prone. Learn the code-frame workflow, the manual vs AI tradeoff, and how Koji auto-codes verbatims and captures the depth a survey verbatim never could.