{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-06T14:12:46.945Z"},"content":[{"type":"documentation","id":"1841acc4-3322-4b57-9cda-1c5181723ac0","slug":"keyness-analysis-distinctive-segment-language","title":"Keyness: Which Words Are Actually Distinctive to a Customer Segment (2026)","url":"https://www.koji.so/docs/keyness-analysis-distinctive-segment-language","summary":"Keyness compares a target corpus against a reference corpus to find over-represented terms, typically with log-likelihood. Two things decide whether the output is usable. First, the reference corpus determines the answer: compare customers against customers, keep target and reference disjoint, and always report which reference was used. Second, standard keyness treats the corpus as a homogeneous whole, so one verbose transcript can dominate; Egbert and Biber 2019 propose text dispersion keyness, which counts each term once per document. A keyword is a pointer to passages, not a finding.","content":"**Bottom line up front:** If you compare what churned customers said against what retained customers said by counting words, you will rediscover that both groups say \"the\", \"product\" and \"team\" a lot. The statistic you actually want is **keyness**: how much more characteristic a term is of one group than of a reference group. And there is a well-documented trap inside it. Standard keyness is computed over a corpus treated as one undifferentiated block, so it rewards terms that a few long, verbose transcripts happened to repeat. Compute keyness on the number of **documents** containing a term, not the number of occurrences, and the list stops being an artefact of your three chattiest participants.\n\n## What keyness is, in one paragraph\n\nKeyness compares two corpora: a **target** corpus (the group you care about, say the 40 interviews with customers who churned) and a **reference** corpus (the comparison group, say the 180 interviews with customers who renewed). For each term, you test whether its relative frequency in the target is higher than you would expect given its relative frequency in the reference. The conventional test statistic is log-likelihood, popularised for this family of problems by Ted Dunning's 1993 work on the statistics of surprise and coincidence in *Computational Linguistics*; mutual information, introduced to lexicography by Church and Hanks at the 27th Annual Meeting of the Association for Computational Linguistics, is the other long-standing option.\n\nThe output is a ranked list of terms that are over-represented in the target. That list is the closest thing text analysis has to an answer to \"what is different about these people\".\n\n## The reference corpus decides the answer\n\nThis is the part that catches everyone, and it has no statistical fix. Keyness is not a property of your target corpus. It is a property of the **pair**. Change the reference and the keywords change.\n\nRun the same 40 churned-customer interviews against three different references:\n\n| Reference corpus | What rises to the top | What it tells you |\n|---|---|---|\n| Retained customers | pricing, approval, budget | Why these customers left rather than stayed |\n| All customers, churned included | mild versions of the same terms, weaker signal | Diluted, because the target is inside the reference |\n| General English | product, dashboard, onboarding | That you work in software. Useless |\n\nThe third row is the common own-goal: comparing customer interviews against a general-language baseline surfaces your own product vocabulary, which you already knew. The second row is the subtler one. **If your target is a subset of your reference, you are comparing a group against itself plus others, which shrinks every difference toward zero.** Keep the two corpora disjoint.\n\nState your reference corpus in every keyness result you circulate. A keyword list without its reference is uninterpretable, in the same way that a percentage without a denominator is uninterpretable.\n\n## The dispersion trap inside keyness\n\nHere is the failure that the corpus linguistics literature has documented and that almost no feedback tool accounts for.\n\nKeyword analysis has become a standard tool for identifying the words especially characteristic of texts in a target domain. But as Jesse Egbert and Douglas Biber set out in a 2019 study in *Corpora*, the statistical computation of keyness makes no reference to those texts. Once the corpus has been constructed, it is treated as a homogeneous whole for the purpose of computing keyness. The consequence they identify is precise: the keywords in such lists are relatively frequent in the corpus, but they are often not widely dispersed across the texts of that corpus, and are therefore not truly representative of the target discourse domain.\n\nIn feedback terms: one customer who mentions \"latency\" thirty times across a long interview can push \"latency\" to the top of your churn keyword list, even though no other churned customer raised it at all. The statistic cannot see that the thirty occurrences came from one person, because the corpus was flattened before the test ran.\n\nEgbert and Biber's proposed remedy is **text dispersion keyness**, which computes keyness from the number of texts in which a term appears rather than from its corpus frequency. They compared it against four other methods for computing keyness in a series of case studies identifying the keywords typical of online travel blogs, assessing each method on content-generalisability and content-distinctiveness, and concluded that text dispersion keyness is the superior measure for generating keyword lists.\n\nThe practitioner version is almost embarrassingly simple:\n\n> **Count each term at most once per document.** Then run your keyness test on those document counts.\n\nA term now scores highly only if *many different* interviews contain it. Repetition within one interview stops counting. You can implement this with a single change to how you build the input table, and it removes an entire class of false keyword.\n\n## A worked comparison\n\nForty churned interviews against 180 retained interviews. The same corpus, scored two ways.\n\n| Term | Occurrences (churn) | Interviews containing it | Rank by occurrence | Rank by document count |\n|---|---|---|---|---|\n| pricing | 61 | 28 of 40 | 2 | 1 |\n| latency | 34 | 3 of 40 | 3 | 19 |\n| approval | 44 | 22 of 40 | 4 | 2 |\n| migration | 77 | 5 of 40 | 1 | 14 |\n\nRanked by occurrence, your churn story is about migration and latency. Ranked by document count, it is about pricing and approval. The second story is the one that generalises: migration dominated the raw count because two customers described a migration at length, and a theme present in 5 of 40 interviews is not what characterises the group.\n\nBoth columns are worth keeping. A high-occurrence, low-document term like migration is a genuine signal about a small number of accounts, and it belongs in your notes. It simply does not belong at the top of a list labelled \"what churned customers talk about\".\n\n## Choosing and reading a keyness list\n\nA workable procedure:\n\n1. **Define the two corpora and keep them disjoint.** Target and reference must not overlap.\n2. **Normalise the comparison for size.** The groups will differ in interview count and length; the test statistic handles this, but only if you feed it correct totals.\n3. **Count once per document.** Per the dispersion fix above.\n4. **Strip your own product vocabulary.** Your brand name, feature names and the words from your own question wording will dominate otherwise. Participants echo the question they were asked.\n5. **Read the top 30 in context before interpreting any of them.** A keyword is a pointer to passages, not a finding. \"Approval\" could mean procurement friction or a praised approvals feature; only the surrounding text distinguishes them.\n6. **Report the reference corpus alongside the list.** Always.\n\nStep 4 deserves emphasis because it interacts with how you collect data. If your interview guide asks about pricing, \"pricing\" will be key in every corpus you build, and you will have discovered your own script.\n\n## How Koji handles this\n\nKeyness analysis needs two things that are hard to get from a pile of transcripts: clean group membership, and the ability to count per document rather than per occurrence. Koji provides both.\n\n- **Clean group labels come from structured questions.** Koji supports six first-class structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - so plan, segment, tenure and renewal intent arrive as structured values rather than something you infer from prose. That is what makes a disjoint target and reference corpus definable in the first place. The [structured questions guide](/docs/structured-questions-guide) covers how each type is asked and aggregated.\n- **Stable question IDs keep the document boundary intact.** Every Koji study question carries a stable identifier that preserves traceability from the interview plan through the AI interviewer to analysis and report aggregation. Because each response stays attached to its participant and its question, counting a term once per interview, or once per account, is a grouping choice rather than a data-recovery project.\n- **Comparable instruments across groups.** A keyness comparison is only valid if both groups were asked comparable questions. Koji's AI interviewer works from the same study definition for every participant, so your churned and retained corpora differ because the customers differ, not because a human moderator asked the churn cohort a pointed extra question.\n- **Insights chat for the context step.** Step 5 above, reading keywords in context, is the step teams skip. Koji lets you query the corpus conversationally and pull the passages behind a term, so checking whether \"approval\" means friction or praise takes one question instead of a transcript search.\n- **Quality scores to screen the input.** Koji scores each interview 1-5 across relevance, depth and coverage. Thin interviews contribute few terms and distort document counts; filtering them before a keyness run is a two-second change in Koji and a meaningful improvement in the output.\n\nUsed this way, Koji turns a keyness analysis from a scripting exercise into a filter-and-read workflow, which matters because the value of keyness is almost entirely in the reading.\n\n## Common mistakes\n\n- **Comparing against general English.** Surfaces your own product vocabulary. Compare customers against customers.\n- **Letting the target sit inside the reference.** Dilutes every difference. Keep them disjoint.\n- **Ranking by raw occurrence.** Lets one verbose participant write your headline, which is the Egbert and Biber objection.\n- **Treating a keyword as a finding.** It is a pointer to passages. Read them.\n- **Forgetting your question wording is in the corpus.** You will find the words you asked about.\n- **Circulating a list without its reference corpus.** Nobody downstream can interpret it.\n\n## Frequently asked questions\n\n### What is keyness analysis?\n\nKeyness analysis identifies the terms that are over-represented in one corpus relative to a reference corpus, using a statistical test such as log-likelihood. In customer research it answers questions like \"what do churned customers talk about that retained customers do not\". The output is a ranked list of distinctive terms rather than common ones, which is what distinguishes it from a word frequency count or a word cloud.\n\n### How is keyness different from just counting word frequency?\n\nWord frequency tells you what is common; keyness tells you what is distinctive. Frequency lists from any two customer groups look nearly identical, because both are dominated by ordinary English and your own product vocabulary. Keyness divides out that shared baseline by comparing against a reference corpus, so what remains is the difference between the groups rather than the content they share.\n\n### Which reference corpus should I use?\n\nUse the most similar group that excludes your target. To characterise churned customers, compare them against retained customers, not against all customers and not against general English. Comparing against general English returns your product vocabulary, and comparing against a superset that contains your target shrinks every difference toward zero. Whatever you choose, report it with the results, because a keyword list cannot be interpreted without knowing what it was compared against.\n\n### Why do my keyword lists keep surfacing things only one customer said?\n\nBecause standard keyness is computed over the corpus as a single undifferentiated block, so thirty repetitions by one talkative participant count the same as thirty mentions spread across thirty people. Egbert and Biber documented exactly this, noting that conventional keywords are frequent in the corpus but often not widely dispersed across its texts. The fix is to count each term at most once per interview and run the test on those document counts.\n\n### What is text dispersion keyness?\n\nText dispersion keyness computes keyness from the number of texts in which a term appears rather than from its total frequency in the corpus. Egbert and Biber proposed it in 2019, compared it against four other keyness methods across case studies on online travel blogs, and found it superior for generating keyword lists when judged on content-generalisability and content-distinctiveness. In practice it means counting once per document, which makes a term score highly only when many different documents contain it.\n\n### Can I run keyness analysis on Koji interview data?\n\nYes, and the two prerequisites are what Koji is built to supply. Group membership comes from Koji's structured question types, so you can define disjoint target and reference corpora from structured values rather than inferring them from prose. Stable question IDs keep every response attached to its participant and question, so you can count terms once per interview or once per account as needed. Koji's insights chat then covers the step that matters most, which is reading the passages behind each keyword before you interpret it.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types that give you the clean group labels a keyness comparison requires\n- [Theme Dispersion vs Mention Frequency](/docs/theme-dispersion-vs-mention-frequency) - the same dispersion problem applied to coded themes rather than terms\n- [Verbatim Analysis Guide](/docs/verbatim-analysis-guide) - coding open-ended responses at scale, the step before any keyness run\n- [Measurement Invariance When Comparing Groups](/docs/measurement-invariance-comparing-groups) - whether your two groups were measured comparably in the first place\n- [Quotes in Context](/docs/archival-bond-research-quotes-context) - why a keyword divorced from its passage loses its meaning\n- [Topic Modeling for Customer Feedback](/docs/topic-modeling-customer-feedback) - the unsupervised alternative when you have no groups to compare\n","category":"Analysis & Synthesis","lastModified":"2026-10-06T07:51:48.241748+00:00","metaTitle":"Keyness Analysis for Customer Segments (2026)","metaDescription":"Keyness finds terms distinctive to one customer group. Use a disjoint reference corpus and count once per document, not per occurrence.","keywords":["keyness analysis","distinctive terms customer segment","what words do churned users use","compare feedback between segments","text dispersion keyness","log-likelihood keyness"],"aiSummary":"Keyness compares a target corpus against a reference corpus to find over-represented terms, typically with log-likelihood. Two things decide whether the output is usable. First, the reference corpus determines the answer: compare customers against customers, keep target and reference disjoint, and always report which reference was used. Second, standard keyness treats the corpus as a homogeneous whole, so one verbose transcript can dominate; Egbert and Biber 2019 propose text dispersion keyness, which counts each term once per document. A keyword is a pointer to passages, not a finding.","aiPrerequisites":["A corpus of open-ended responses split into two comparable groups","Group labels such as churned vs retained"],"aiLearningOutcomes":["Explain how keyness differs from raw word frequency","Select a disjoint and appropriate reference corpus","Apply text dispersion keyness by counting once per document","Avoid surfacing your own product vocabulary and question wording"],"aiDifficulty":"intermediate","aiEstimatedTime":"10 min"}],"pagination":{"total":1,"returned":1,"offset":0}}