{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-02T09:38:41.437Z"},"content":[{"type":"documentation","id":"78313930-fb30-45ba-a9ec-cb430147a98b","slug":"open-end-lexical-diversity-diagnostic","title":"Has Your Open-End Data Gone Flat? Measuring Lexical Diversity Instead of Accusing Participants","url":"https://www.koji.so/docs/open-end-lexical-diversity-diagnostic","summary":"A batch-level diagnostic for detecting homogenization in open-ended research data without making individual-level accusations. Run three numbers per question per wave: MTLD, near-duplicate rate, and median word count. The diagnostic fingerprint of generic text is answers getting longer while diversity falls; shorter plus less varied indicates fatigue instead, which has the opposite remedy. Raw type-token ratio must not be compared across batches of differing length because it falls mechanically with length -- a 50-word response with 40 types scores 0.80 against a 200-word response with 110 types at 0.55, inverting the true ranking. McCarthy and Jarvis (Behavior Research Methods, 2010, 42(2):381-392) found MTLD the only index not varying as a function of text length. Compare a question only against its own history; never use the statistic against a participant.","content":"You cannot responsibly decide whether any single open-ended response was written by a language model. You can measure whether a whole batch of them has lost its variety. Lexical diversity and near-duplicate rate, tracked wave over wave on the same question, give you a defensible early-warning signal that accuses nobody and has no disparate impact on any group of participants.\n\n## The short answer\n\nRun three numbers on every wave of open-ended responses, per question:\n\n| Measure | What it answers | Watch for |\n| --- | --- | --- |\n| MTLD | How varied is the vocabulary, independent of answer length | A fall against the same question's prior waves |\n| Near-duplicate rate | What share of answers are near-copies of each other | Any rise, especially above a few percent |\n| Median word count | Are answers getting longer | A rise alongside falling diversity |\n\nThe combination is what carries the signal. Answers getting **longer while becoming less varied** is the specific fingerprint of generic text entering your corpus. Either movement alone has innocent explanations; together they rarely do.\n\nCrucially, this is a statement about a batch, not about a person. That is what makes it usable. It is the defensible counterpart to the individual-level detection that [AI detectors cannot deliver](/docs/ai-detector-false-positives-research).\n\n## Why a corpus statistic is the right altitude\n\nThe reason to move up a level is not squeamishness. It is that the individual-level question has no reliable answer and the batch-level question does.\n\nConsider what you actually need to know. You do not need to know that participant 207 used a chatbot. You need to know whether the themes you are about to present to your product team still reflect the range of what your customers think. That is a property of the corpus. Asking it at the corpus level is both answerable and the question you genuinely had.\n\nThis also resolves the ethical problem cleanly. A corpus statistic triggers an action against an *instrument* -- rewrite the question, shorten the study, switch to voice -- rather than against a participant. Nobody is denied payment, nobody is accused, and no group gets filtered out of your sample. You can publish the method in your write-up without anyone objecting to it.\n\n## The three measures worth running\n\n### Type-token ratio, and why not to trust it raw\n\nType-token ratio, or TTR, is unique words divided by total words. It is the obvious first thing to compute, and on its own it will mislead you here -- specifically and badly.\n\nTTR falls mechanically as text gets longer, because common words repeat. Work an example:\n\n- Response A: 50 words, 40 of them distinct. TTR = 40 / 50 = 0.80\n- Response B: 200 words, 110 of them distinct. TTR = 110 / 200 = 0.55\n\nResponse B has nearly three times the distinct vocabulary of response A, and a TTR that is far lower. Raw TTR calls the richer answer the poorer one.\n\nNow notice why this matters so much for this particular problem. AI-assisted answers tend to be **both longer and more homogeneous**. Raw TTR confounds the two effects: it will drop partly because the vocabulary genuinely narrowed and partly just because answers got longer. You cannot tell from the number how much of the fall is real. You will \"detect\" homogenization in a batch where answers merely got wordier, and you will under-read it elsewhere.\n\nSo compute TTR if you like, but never compare TTR across batches whose answer lengths differ.\n\n### MTLD, the length-invariant choice\n\nThe measure built to solve exactly this is MTLD, the measure of textual lexical diversity. The validation study is McCarthy and Jarvis, \"MTLD, vocd-D, and HD-D: a validation study of sophisticated approaches to lexical diversity assessment,\" in *Behavior Research Methods*, 2010, volume 42, issue 2, pages 381 to 392.\n\nTheir conclusion is the reason to prefer it: MTLD \"performs well with respect to all four types of validity and is, in fact, the only index not found to vary as a function of text length.\"\n\nThat property is the whole point. It lets you compare this wave against last wave even though the answers got longer, which is precisely the comparison the length confound would otherwise destroy. The same paper reports that HD-D is a viable alternative to the established vocd-D measure, and that MTLD, vocd-D or HD-D, and Maas each appear to capture somewhat different lexical information -- so if you have the means, reporting more than one is better than relying on a single index.\n\nPractically: compute MTLD per response, then take the median across the batch for a given question. Track that median over time. You are looking for a step change, not an absolute threshold -- there is no universal \"good\" MTLD value, because it depends on your question, your audience and your domain vocabulary.\n\n### Near-duplicate rate\n\nLexical diversity measures variety *within* a response. Near-duplicate rate measures variety *between* responses, and it catches a failure mode diversity scores miss: twenty answers that are each individually rich but all say the same thing in the same shape.\n\nCompare every pair of responses to a question on token overlap and count the share of responses that have at least one near-twin above some similarity cut. The absolute value depends on your cut, so again, track the trend rather than the level. A healthy open-ended question on a varied population produces very few near-twins. A question that has started attracting generic answers produces clusters, because generic text converges.\n\nThis measure has a useful side benefit: it also catches ordinary copy-paste duplication and template answers from a single participant across a repeated study, which are older problems than AI.\n\n## Reading the signal\n\nThe three numbers combine into a small decision table:\n\n| MTLD | Near-duplicates | Length | Most likely reading |\n| --- | --- | --- | --- |\n| Stable | Stable | Stable | Healthy. Do nothing. |\n| Falling | Rising | Rising | Generic text entering the corpus. Investigate the question. |\n| Falling | Stable | Falling | Fatigue or disengagement, not outsourcing. Shorten the instrument. |\n| Stable | Rising | Stable | Possible template or shared-answer behaviour, or a genuinely converged view. |\n| Rising | Stable | Rising | Usually good. More engaged participants giving fuller answers. |\n\nThe third row is worth dwelling on, because it is the one teams misread most often. Short, repetitive, low-diversity answers are the signature of [survey fatigue](/docs/survey-fatigue) and low effort, not of chatbot use -- and the remedy is the opposite one. Outsourcing makes answers longer; exhaustion makes them shorter. Checking length tells you which problem you have, and it is the cheapest of the three numbers to compute.\n\nA fall in diversity is a prompt to look at your instrument, in roughly this order: Is this question answerable without the participant's specific experience? Is it late in a long study? Does it demand more writing than it is worth? Has the recruitment source changed?\n\n## What this is not\n\n- **Not a verdict on any response.** Nothing here licenses dropping an individual answer. If you want to exclude responses you need a documented rule justified on other grounds, reported with its effect on sample size.\n- **Not a measure of insight quality.** A batch can have excellent lexical diversity and tell you nothing useful, and a tightly-worded batch on a narrow technical question can be genuinely informative at low diversity.\n- **Not comparable across questions.** Diversity depends heavily on what you asked. Compare a question against its own history, never against a different question.\n- **Not a substitute for reading the data.** These numbers tell you where to look. The looking is still your job, and [thematic analysis](/docs/thematic-analysis-guide) is still how the findings get made.\n\n## How Koji handles this\n\nKoji is built so that the corpus stays comparable over time, which is the precondition for any of this working.\n\n- **Stable question IDs make wave-over-wave comparison valid.** Koji carries a stable identifier for each question from the research brief through the AI interviewer to analysis and report aggregation. You are therefore comparing the same question to itself across waves, rather than hoping two similarly-worded questions are equivalent.\n- **Adaptive probing raises genuine diversity rather than faking it.** Because Koji's AI interviewer generates each follow-up from what the participant just said, no two transcripts follow the same path. The variety is real, and it comes from the participant's own content.\n- **Coding distinguishes the participant's framing from the analyst's.** Koji labels open-ended codes as either descriptive or in vivo, so framings that capture a participant's specific wording are marked as such. A corpus losing its in vivo codes is losing exactly the distinctive material you were measuring for.\n- **Structured questions keep open-ends short and focused.** With six question types -- `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` -- the counting moves into typed fields. Short, specific open-ends have less length confound and far less outsourcing pressure.\n- **Quality scoring is a complementary per-interview read.** Koji scores each interview 1-5 across relevance, depth and coverage. Diversity is your batch-level signal; the depth score is a per-interview one, and the two disagreeing is itself informative.\n\n## Common mistakes\n\n- **Comparing raw TTR across batches with different answer lengths.** This is the single most common error, and the worked example above shows it inverting the ranking. Use MTLD, or compare only within a narrow length band.\n- **Setting an absolute diversity threshold.** There is no universal cut. The signal is movement against the same question's own history.\n- **Acting on one wave.** You need a baseline. Start computing these now even if you have no concern, so that you have history when you do.\n- **Treating a fall in diversity as proof of AI use.** It is consistent with fatigue, a changed recruitment source, a narrower question, or a genuinely converged customer view. Check length and near-duplicates before concluding anything.\n- **Using the statistic against a participant.** It does not support individual conclusions, and using it that way reintroduces every fairness problem that made detectors unusable.\n\n## Frequently asked questions\n\n### What is lexical diversity and why would I measure it on survey data?\n\nLexical diversity is how varied the vocabulary in a text is. Measured across a batch of open-ended responses and tracked over time, it tells you whether your answers are still carrying the range of expression they used to. That matters because open-ended questions exist to surface variation, so a corpus losing its variety is losing the thing you collected it for, even when every individual answer still reads well.\n\n### Why should I use MTLD instead of a simple type-token ratio?\n\nBecause type-token ratio falls mechanically as texts get longer, and the problem you are investigating tends to make answers longer. McCarthy and Jarvis found MTLD to be the only index in their validation study not found to vary as a function of text length, which is what allows a fair comparison between waves whose answer lengths differ. A raw type-token comparison across different lengths can even rank a richer answer below a poorer one.\n\n### What counts as a worrying drop in diversity?\n\nThere is no universal number, which is why the method is comparative. Compare a question against its own previous waves and look for a step change rather than drift. The reading is much stronger when diversity falls at the same time as median answer length rises and near-duplicate rate rises, because that specific combination is hard to explain by fatigue or by a changed question.\n\n### Can I use this to decide which responses to exclude?\n\nNo. This is a batch-level diagnostic and it does not support conclusions about individual responses. Using it that way would reintroduce the fairness problems that make detector-based exclusion indefensible. If you need an exclusion rule, base it on documented criteria and report how many responses it removed and why.\n\n### Does falling diversity always mean participants are using AI?\n\nNo, and assuming so will send you after the wrong fix. Falling diversity with shorter answers usually means fatigue or disengagement, and the remedy is a shorter instrument. Falling diversity with longer answers is the pattern that points at generic text. It can also reflect a changed recruitment source, a question you narrowed, or a customer view that has genuinely converged.\n\n### How often should I run these numbers?\n\nOnce per wave, per open-ended question, as a standing part of your analysis prep. The cost is near zero once scripted and the value is entirely in the time series, so the main thing is to start early and keep the question wording stable enough that the comparison stays meaningful.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and moving countable load out of open-ends\n- [Why You Cannot Gate Research on an AI Detector](/docs/ai-detector-false-positives-research) - why the individual-level alternative fails\n- [Thematic Analysis Guide](/docs/thematic-analysis-guide) - turning open-ended data into findings\n- [Survey Data Quality](/docs/survey-data-quality-guide) - the wider detection and prevention workflow\n- [Survey Fatigue](/docs/survey-fatigue) - the competing explanation for falling diversity\n- [Inter-Rater Reliability](/docs/inter-rater-reliability-qualitative-research) - why homogeneous text can inflate coder agreement","category":"Analysis & Synthesis","lastModified":"2026-10-02T03:47:16.127357+00:00","metaTitle":"Lexical Diversity Monitoring for Open-Ended Research Data (2026)","metaDescription":"Measure whether a batch of open-ends lost its variance, using MTLD and near-duplicate rate. A diagnostic that accuses nobody.","keywords":["open-end lexical diversity","mtld lexical diversity","type-token ratio survey","response homogenization","near-duplicate responses","open-ended data quality"],"aiSummary":"A batch-level diagnostic for detecting homogenization in open-ended research data without making individual-level accusations. Run three numbers per question per wave: MTLD, near-duplicate rate, and median word count. The diagnostic fingerprint of generic text is answers getting longer while diversity falls; shorter plus less varied indicates fatigue instead, which has the opposite remedy. Raw type-token ratio must not be compared across batches of differing length because it falls mechanically with length -- a 50-word response with 40 types scores 0.80 against a 200-word response with 110 types at 0.55, inverting the true ranking. McCarthy and Jarvis (Behavior Research Methods, 2010, 42(2):381-392) found MTLD the only index not varying as a function of text length. Compare a question only against its own history; never use the statistic against a participant.","aiPrerequisites":["Familiarity with open-ended survey questions","Comfort reading simple ratios and medians"],"aiLearningOutcomes":["Explain why corpus-level measurement is answerable where individual-level detection is not","Compute and interpret type-token ratio and recognise its length confound","Select MTLD for cross-wave comparison and justify the choice","Read the combination of diversity, duplication and length as a decision table","Distinguish homogenization from fatigue using answer length"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}