Has Your Open-End Data Gone Flat? Measuring Lexical Diversity Instead of Accusing Participants
You cannot judge whether one response was AI-written, but you can measure whether a whole batch has lost its variance. A practical guide to lexical diversity and near-duplicate monitoring for open-ended research data.
You cannot responsibly decide whether any single open-ended response was written by a language model. You can measure whether a whole batch of them has lost its variety. Lexical diversity and near-duplicate rate, tracked wave over wave on the same question, give you a defensible early-warning signal that accuses nobody and has no disparate impact on any group of participants.
The short answer
Run three numbers on every wave of open-ended responses, per question:
| Measure | What it answers | Watch for |
|---|---|---|
| MTLD | How varied is the vocabulary, independent of answer length | A fall against the same question's prior waves |
| Near-duplicate rate | What share of answers are near-copies of each other | Any rise, especially above a few percent |
| Median word count | Are answers getting longer | A rise alongside falling diversity |
The combination is what carries the signal. Answers getting longer while becoming less varied is the specific fingerprint of generic text entering your corpus. Either movement alone has innocent explanations; together they rarely do.
Crucially, this is a statement about a batch, not about a person. That is what makes it usable. It is the defensible counterpart to the individual-level detection that AI detectors cannot deliver.
Why a corpus statistic is the right altitude
The reason to move up a level is not squeamishness. It is that the individual-level question has no reliable answer and the batch-level question does.
Consider what you actually need to know. You do not need to know that participant 207 used a chatbot. You need to know whether the themes you are about to present to your product team still reflect the range of what your customers think. That is a property of the corpus. Asking it at the corpus level is both answerable and the question you genuinely had.
This also resolves the ethical problem cleanly. A corpus statistic triggers an action against an instrument -- rewrite the question, shorten the study, switch to voice -- rather than against a participant. Nobody is denied payment, nobody is accused, and no group gets filtered out of your sample. You can publish the method in your write-up without anyone objecting to it.
The three measures worth running
Type-token ratio, and why not to trust it raw
Type-token ratio, or TTR, is unique words divided by total words. It is the obvious first thing to compute, and on its own it will mislead you here -- specifically and badly.
TTR falls mechanically as text gets longer, because common words repeat. Work an example:
- Response A: 50 words, 40 of them distinct. TTR = 40 / 50 = 0.80
- Response B: 200 words, 110 of them distinct. TTR = 110 / 200 = 0.55
Response B has nearly three times the distinct vocabulary of response A, and a TTR that is far lower. Raw TTR calls the richer answer the poorer one.
Now notice why this matters so much for this particular problem. AI-assisted answers tend to be both longer and more homogeneous. Raw TTR confounds the two effects: it will drop partly because the vocabulary genuinely narrowed and partly just because answers got longer. You cannot tell from the number how much of the fall is real. You will "detect" homogenization in a batch where answers merely got wordier, and you will under-read it elsewhere.
So compute TTR if you like, but never compare TTR across batches whose answer lengths differ.
MTLD, the length-invariant choice
The measure built to solve exactly this is MTLD, the measure of textual lexical diversity. The validation study is McCarthy and Jarvis, "MTLD, vocd-D, and HD-D: a validation study of sophisticated approaches to lexical diversity assessment," in Behavior Research Methods, 2010, volume 42, issue 2, pages 381 to 392.
Their conclusion is the reason to prefer it: MTLD "performs well with respect to all four types of validity and is, in fact, the only index not found to vary as a function of text length."
That property is the whole point. It lets you compare this wave against last wave even though the answers got longer, which is precisely the comparison the length confound would otherwise destroy. The same paper reports that HD-D is a viable alternative to the established vocd-D measure, and that MTLD, vocd-D or HD-D, and Maas each appear to capture somewhat different lexical information -- so if you have the means, reporting more than one is better than relying on a single index.
Practically: compute MTLD per response, then take the median across the batch for a given question. Track that median over time. You are looking for a step change, not an absolute threshold -- there is no universal "good" MTLD value, because it depends on your question, your audience and your domain vocabulary.
Near-duplicate rate
Lexical diversity measures variety within a response. Near-duplicate rate measures variety between responses, and it catches a failure mode diversity scores miss: twenty answers that are each individually rich but all say the same thing in the same shape.
Compare every pair of responses to a question on token overlap and count the share of responses that have at least one near-twin above some similarity cut. The absolute value depends on your cut, so again, track the trend rather than the level. A healthy open-ended question on a varied population produces very few near-twins. A question that has started attracting generic answers produces clusters, because generic text converges.
This measure has a useful side benefit: it also catches ordinary copy-paste duplication and template answers from a single participant across a repeated study, which are older problems than AI.
Reading the signal
The three numbers combine into a small decision table:
| MTLD | Near-duplicates | Length | Most likely reading |
|---|---|---|---|
| Stable | Stable | Stable | Healthy. Do nothing. |
| Falling | Rising | Rising | Generic text entering the corpus. Investigate the question. |
| Falling | Stable | Falling | Fatigue or disengagement, not outsourcing. Shorten the instrument. |
| Stable | Rising | Stable | Possible template or shared-answer behaviour, or a genuinely converged view. |
| Rising | Stable | Rising | Usually good. More engaged participants giving fuller answers. |
The third row is worth dwelling on, because it is the one teams misread most often. Short, repetitive, low-diversity answers are the signature of survey fatigue and low effort, not of chatbot use -- and the remedy is the opposite one. Outsourcing makes answers longer; exhaustion makes them shorter. Checking length tells you which problem you have, and it is the cheapest of the three numbers to compute.
A fall in diversity is a prompt to look at your instrument, in roughly this order: Is this question answerable without the participant's specific experience? Is it late in a long study? Does it demand more writing than it is worth? Has the recruitment source changed?
What this is not
- Not a verdict on any response. Nothing here licenses dropping an individual answer. If you want to exclude responses you need a documented rule justified on other grounds, reported with its effect on sample size.
- Not a measure of insight quality. A batch can have excellent lexical diversity and tell you nothing useful, and a tightly-worded batch on a narrow technical question can be genuinely informative at low diversity.
- Not comparable across questions. Diversity depends heavily on what you asked. Compare a question against its own history, never against a different question.
- Not a substitute for reading the data. These numbers tell you where to look. The looking is still your job, and thematic analysis is still how the findings get made.
How Koji handles this
Koji is built so that the corpus stays comparable over time, which is the precondition for any of this working.
- Stable question IDs make wave-over-wave comparison valid. Koji carries a stable identifier for each question from the research brief through the AI interviewer to analysis and report aggregation. You are therefore comparing the same question to itself across waves, rather than hoping two similarly-worded questions are equivalent.
- Adaptive probing raises genuine diversity rather than faking it. Because Koji's AI interviewer generates each follow-up from what the participant just said, no two transcripts follow the same path. The variety is real, and it comes from the participant's own content.
- Coding distinguishes the participant's framing from the analyst's. Koji labels open-ended codes as either descriptive or in vivo, so framings that capture a participant's specific wording are marked as such. A corpus losing its in vivo codes is losing exactly the distinctive material you were measuring for.
- Structured questions keep open-ends short and focused. With six question types --
open_ended,scale,single_choice,multiple_choice,rankingandyes_no-- the counting moves into typed fields. Short, specific open-ends have less length confound and far less outsourcing pressure. - Quality scoring is a complementary per-interview read. Koji scores each interview 1-5 across relevance, depth and coverage. Diversity is your batch-level signal; the depth score is a per-interview one, and the two disagreeing is itself informative.
Common mistakes
- Comparing raw TTR across batches with different answer lengths. This is the single most common error, and the worked example above shows it inverting the ranking. Use MTLD, or compare only within a narrow length band.
- Setting an absolute diversity threshold. There is no universal cut. The signal is movement against the same question's own history.
- Acting on one wave. You need a baseline. Start computing these now even if you have no concern, so that you have history when you do.
- Treating a fall in diversity as proof of AI use. It is consistent with fatigue, a changed recruitment source, a narrower question, or a genuinely converged customer view. Check length and near-duplicates before concluding anything.
- Using the statistic against a participant. It does not support individual conclusions, and using it that way reintroduces every fairness problem that made detectors unusable.
Frequently asked questions
What is lexical diversity and why would I measure it on survey data?
Lexical diversity is how varied the vocabulary in a text is. Measured across a batch of open-ended responses and tracked over time, it tells you whether your answers are still carrying the range of expression they used to. That matters because open-ended questions exist to surface variation, so a corpus losing its variety is losing the thing you collected it for, even when every individual answer still reads well.
Why should I use MTLD instead of a simple type-token ratio?
Because type-token ratio falls mechanically as texts get longer, and the problem you are investigating tends to make answers longer. McCarthy and Jarvis found MTLD to be the only index in their validation study not found to vary as a function of text length, which is what allows a fair comparison between waves whose answer lengths differ. A raw type-token comparison across different lengths can even rank a richer answer below a poorer one.
What counts as a worrying drop in diversity?
There is no universal number, which is why the method is comparative. Compare a question against its own previous waves and look for a step change rather than drift. The reading is much stronger when diversity falls at the same time as median answer length rises and near-duplicate rate rises, because that specific combination is hard to explain by fatigue or by a changed question.
Can I use this to decide which responses to exclude?
No. This is a batch-level diagnostic and it does not support conclusions about individual responses. Using it that way would reintroduce the fairness problems that make detector-based exclusion indefensible. If you need an exclusion rule, base it on documented criteria and report how many responses it removed and why.
Does falling diversity always mean participants are using AI?
No, and assuming so will send you after the wrong fix. Falling diversity with shorter answers usually means fatigue or disengagement, and the remedy is a shorter instrument. Falling diversity with longer answers is the pattern that points at generic text. It can also reflect a changed recruitment source, a question you narrowed, or a customer view that has genuinely converged.
How often should I run these numbers?
Once per wave, per open-ended question, as a standing part of your analysis prep. The cost is near zero once scripted and the value is entirely in the time series, so the main thing is to start early and keep the question wording stable enough that the comparison stays meaningful.
Related Resources
- Structured Questions Guide - the six question types, and moving countable load out of open-ends
- Why You Cannot Gate Research on an AI Detector - why the individual-level alternative fails
- Thematic Analysis Guide - turning open-ended data into findings
- Survey Data Quality - the wider detection and prevention workflow
- Survey Fatigue - the competing explanation for falling diversity
- Inter-Rater Reliability - why homogeneous text can inflate coder agreement
Related Articles
Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement
Learn how to measure inter-rater (intercoder) reliability in qualitative research using Cohen's kappa and Krippendorff's alpha, what thresholds count as reliable, and how AI-native tools make consistent coding the default.
Partial Interviews: Should You Analyse Someone Who Answered Half Your Questions?
A partial interview is breakoff - a third category that is neither unit nonresponse nor item nonresponse. How Koji flags partials, why they usually cost you nothing, and when to include them.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Data Quality: How to Detect and Prevent Bad Responses (2026)
The threats that corrupt survey data — straightlining, speeding, bots, fraud, and inattentive respondents — how to detect and prevent each, and why conversational AI interviews are structurally resistant to the junk that plagues panel surveys.
Survey Fatigue: Why It's Getting Worse (And How AI Interviews Solve It)
Survey fatigue is driving response rates to historic lows. This guide explains why it is happening, what it costs your research, and how AI-moderated interviews deliver better data without burning out respondents.
The Complete Guide to Thematic Analysis
Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.