AI-Assisted Answers: What to Do When Participants Let an LLM Write Their Open-Ends
A real, eligible participant pasting your open-ended question into a chatbot is a validity problem, not a fraud problem. Here is what the evidence shows, why homogenization is the expensive part, and how to design questions a generic answer cannot survive.
When a participant pastes your open-ended question into a chatbot and pastes the answer back, you have a validity problem, not a fraud problem. The person is real, eligible and willing. Only the words are not theirs. That distinction decides your entire response: detection and exclusion is the wrong instinct, and question design plus corpus-level monitoring is the right one.
The short answer
Treat AI-assisted open-ends as a measurement failure you caused, not a betrayal the participant committed. Three moves, in order of value:
- Design questions a generic answer cannot answer. Ask for a specific episode with a date, a sequence, or a number the participant alone could know. A language model can produce a fluent paragraph about onboarding friction; it cannot tell you what happened the Tuesday your billing page timed out.
- Monitor the batch, not the person. You cannot responsibly judge any single response, but you can measure whether a whole wave of open-ends has lost its variance. That is a corpus statistic, and it is defensible.
- Stop rewarding length. Most outsourcing is a rational response to a question that demanded an essay for a flat incentive.
What you should not do is run responses through an AI detector and drop the flagged ones. The false-positive behaviour of those tools makes that indefensible, and it punishes exactly the participants you can least afford to lose.
This is not fraud, and the difference is load-bearing
Research operations already has vocabulary for bad responses, and AI-assisted answers fit none of the existing boxes:
| Failure | Who is answering | What is wrong | Existing playbook |
|---|---|---|---|
| Survey fraud | A bot, or a person misrepresenting eligibility | The respondent is not who they claim | Identity and eligibility checks |
| Satisficing or inattention | A real, eligible person | Effort is too low to be informative | Attention checks, shorter instruments |
| Synthetic users | Nobody. The researcher substituted a model for people | There is no respondent at all | A methodology decision, made by you |
| AI-assisted open-ends | A real, eligible, willing person | The person is genuine; the prose is not theirs | Mostly missing |
The practical consequence: every control built for the first three rows misses the fourth. An eligibility screener passes this participant, because they are eligible. An attention check passes them, because they are paying attention -- arguably more attention than someone typing six words. A fraud scorecard that looks for implausible speed can actively mislead you, because a participant who stops to prompt a chatbot may take longer, not less time, than an honest fast typist.
So the gap is real, and it is not covered by survey fraud detection, attention checks, or the synthetic users debate, all of which answer different questions.
What the evidence actually shows
The best current measurement comes from Zhang, Xu and Alvero, published in Sociological Methods & Research in 2025 under the title "Generative AI Meets Open-Ended Survey Responses: Research Participant Use of AI and Homogenization."
Two findings matter, and they come from two different parts of the study, which is worth keeping straight.
The prevalence figure. In an original survey of research participants recruited from a popular online platform for sourcing social science research subjects, 34 percent reported using LLMs to help them answer open-ended survey questions.
Read that population description carefully before you reuse the number. These are paid participants on a research-subject platform -- people for whom answering surveys is a repeated, compensated activity. That is the highest-exposure population there is. It is not a measurement of your own customers, and quoting "34 percent of customers use AI" would be a straightforward misreading. Use it as evidence that the behaviour is common and normalised in paid panels, and as a reason to design defensively everywhere else.
The homogenization finding. This did not come from watching those participants. The authors ran simulations comparing human-written responses from three pre-ChatGPT studies against LLM-generated text, and found that LLM responses are more homogeneous and positive, particularly when they describe social groups in sensitive questions. Their stated concern is that these patterns may mask important underlying social variation in attitudes and beliefs among human subjects.
That is the sentence to tape to your monitor, because it names the real cost.
Why homogenization is the expensive part
Most teams worry about AI-assisted answers being fake. The more serious problem is that they are average.
Qualitative research does not exist to find the modal opinion. If you wanted the modal opinion you would run a structured scale question and count. Open-ends exist to surface the thing you did not think to ask about: the unusual workaround, the specific trigger, the minority use case that turns out to be a segment. That signal lives in the tail of the distribution.
A language model is, by construction, a machine for producing the centre of a distribution. It is trained and tuned to emit the fluent, agreeable, broadly representative answer. So when a share of your open-ends come from a model, you do not get noise scattered around the truth -- you get mass pulled toward the middle, and the tail thins out.
This has an unpleasant property: it makes your data look better. Variance drops. Responses get longer and more articulate. Themes consolidate cleanly, and inter-coder agreement may even improve, because homogeneous text is easier to code consistently. Every surface signal of quality improves while the thing you were buying -- variation -- quietly disappears. Compare this with low-effort responses, which announce themselves as junk. AI-assisted responses are camouflaged as excellence.
If you take one operational lesson from this page: a wave of open-ends that suddenly reads better than usual deserves more scrutiny, not less.
The four conditions that invite outsourcing
Participants are not adversaries. They outsource when the instrument makes outsourcing the sensible move. Four design choices do most of the damage:
A question that rewards volume over specificity
"Tell us about your experience with our onboarding process" is an essay prompt. It has no verifiable content, no anchor in time, and no way to be wrong. It is the single most outsourceable question shape there is. Contrast with "What was the last thing that made you stop and ask someone for help?" -- answerable in one sentence, and not answerable at all by a model that has never used your product.
A length cue that implies effort
Minimum character counts, a large empty textarea, and progress copy like "the more detail the better" all signal that length is the currency. Participants who want to comply honestly, and who are tired, will reach for the fastest tool that produces length.
Flat compensation for variable effort
If every open-end pays the same regardless of difficulty, the rational participant minimises cost per unit of acceptable output. This is ordinary satisficing economics operating on a new substrate. The chatbot did not create the incentive; it removed the last friction from acting on it.
An instrument long enough to exhaust goodwill
Outsourcing concentrates late in long instruments, for the same reason survey fatigue and breakoff do. A participant who answered your first four open-ends honestly and is staring at a fifth has a strong motive to finish by other means. The fix is the one you already know: ask for less. See ideal survey length for the trade-offs.
What to do instead
Anchor every open-end in something only this person knows
The strongest defence is epistemic, not technical. Ask for episodes, not opinions:
- "Walk me through the last time you tried to do X. What happened first?"
- "What did you do immediately after that?"
- "Who else was involved, and what did they say?"
- "What did you expect to happen instead?"
A model can invent a plausible answer to each of these, but it cannot keep a specific invented episode internally consistent across three unscripted follow-ups that depend on the previous answer. That is the mechanism that actually works, and it is why conversational formats hold up better than forms here.
Ask about AI use directly, and non-punitively
A plain question near the end -- "Did you use any AI tool to help write your answers today?" with a neutral frame and no penalty attached -- costs you one question and gives you a covariate you can analyse with. It will under-report. It is still far better than a detector score, because it is honest about its own uncertainty and it does not misclassify anyone. Pair it with clear framing up front about why you want their own words. Our guidance on consent and disclosure wording applies directly: the wording is part of the instrument.
Measure the corpus
Track, wave over wave, the lexical diversity and near-duplicate rate of your open-ends. Falling diversity across a batch is a signal you can act on without accusing any individual. This is the defensible counterpart to detection, and it deserves its own treatment.
Never gate on a detector
Covered at length separately, but the short version: the published false-positive rates on non-native English writing are high enough that a detector-based exclusion rule systematically removes non-native speakers from your sample. That is a sampling bias you introduced, and it is worse than the problem you were trying to fix.
How Koji handles this
Koji's architecture attacks the cause rather than the symptom, because an AI-moderated conversation is a fundamentally harder target for outsourcing than a form.
- Adaptive follow-ups break invented episodes. Koji's AI interviewer reads the previous answer and asks the next question from it. A pasted paragraph survives one exchange; it rarely survives being asked what happened next, who else was involved, and what the participant expected instead. The built-in
mom_testframework encodes exactly this, with question patterns like "Walk me through how you currently handle [task]" and "What happened after that?" -- mechanism probes rather than opinion requests. - Voice mode removes the copy-paste surface entirely. In a spoken interview there is no textarea to paste into and no pause long enough to prompt a chatbot without it being obvious. Koji runs both voice and text, so you can choose the mode that matches your risk.
- Structured questions carry the quantitative load, so open-ends can be short. Koji supports six question types --
open_ended,scale,single_choice,multiple_choice,rankingandyes_no. Move the ratings and the choices into typed questions and the open-ends no longer have to carry everything, which removes the length pressure that drives outsourcing in the first place. - Per-interview quality scoring gives you an auditable signal. Koji scores each interview 1-5 with a breakdown across relevance, depth and coverage. A conversation that is fluent but never lands on specifics scores low on depth regardless of word count, which is precisely the failure mode generic prose produces.
- Stable question IDs make the corpus comparable. Because answers map back to specific question IDs from the brief through to report aggregation, you can compare the same question across waves and see diversity move.
Common mistakes
- Treating it as misconduct. It drives the behaviour underground and poisons your disclosure question. The participant thinks they are being helpful.
- Reusing the 34 percent figure for your own customers. It was measured on a paid research-subject platform. Your B2B customers are a different population with different incentives.
- Reading improved writing quality as improved data quality. This is the trap. Fluency went up; variance went down. Only one of those is what you wanted.
- Adding a longer minimum character count. This increases outsourcing pressure. It is the exact wrong lever.
- Excluding suspected responses silently. If you drop responses on suspicion and do not record the rule, your sample is now shaped by an undocumented filter. At minimum, log the rule and report how many responses it removed, as you would for any other exclusion.
Frequently asked questions
Is a participant using AI to answer a survey question cheating?
Generally no, and framing it that way will cost you. The participant is real and eligible, and most people reach for a chatbot because the question asked for more writing than it was worth, not to deceive you. Treat it as a signal that your instrument created the incentive, and fix the question rather than policing the person.
How can I tell whether a specific response was written by an LLM?
Reliably, you cannot, and you should stop trying to decide this at the level of an individual response. There is no test with a false-positive rate low enough to justify acting against one participant. What you can do is measure a whole batch for falling lexical diversity and rising near-duplication, and anchor your questions in specifics so that a generic answer fails to be useful whether or not you can prove its origin.
Should I just ban AI use in my research?
A stated expectation is worth having, because many participants will honour it once they understand you want their own words and why. An enforceable ban is not available to you, so do not build your data-quality plan on one. Combine a clear, non-punitive request with question design that makes generic answers worthless.
Does this affect voice interviews too?
Much less, and that is one of the strongest arguments for voice. A spoken conversation offers no paste target and no unobserved pause in which to prompt a model, and adaptive follow-ups force specifics in real time. Koji runs interviews in both voice and text, so a study with high outsourcing risk can simply be run in voice.
What should I do with responses I already suspect are AI-written?
Do not delete them on suspicion. Keep them, code them, and compare them as a group against the rest of the batch on the dimensions you care about. If they are pulling your themes toward the centre you will see it as reduced variance, which is a finding you can report honestly. Any exclusion rule you do apply should be written down and its effect on sample size reported.
Does a longer, better-written answer mean better data?
No, and assuming so is the most expensive error on this page. Length and fluency are not quality. An answer that is articulate, balanced and generic tells you less than a short, awkward sentence about a specific Tuesday. Judge open-ends by whether they contain something only that person could have said.
Related Resources
- Structured Questions Guide - the six question types, and how to move quantitative load off your open-ends
- Survey Fraud and Respondent Quality - the adjacent problem of fake and ineligible respondents
- Attention Check Questions - catching low-effort responses without annoying real participants
- Synthetic Users in Research - when the researcher, not the participant, substitutes AI for people
- Survey Data Quality - the broader detection and prevention workflow
- Consent Form Wording - why your framing is part of the instrument
Related Articles
Attention Check Questions: How to Catch Low-Effort Survey Responses Without Annoying Real Participants
Attention check questions catch inattentive, low-effort, and fraudulent survey responses. Learn the main types, how many to use, the pitfalls, and why a conversational AI interview reduces the need for them in the first place.
Your Consent Form Is Part of the Instrument (2026)
Consent and confidentiality wording is not administrative text sitting outside your study. A survey experiment shows it changes who finishes - and it filters exactly the people you need on a sensitive topic.
How Long Should a Survey Be? Ideal Survey Length and Question Count
The data-backed guide to ideal survey length — how many questions to ask, how completion rate drops with each question, the 7-minute abandonment cliff, and why conversational AI interviews beat long static surveys.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Data Quality: How to Detect and Prevent Bad Responses (2026)
The threats that corrupt survey data — straightlining, speeding, bots, fraud, and inattentive respondents — how to detect and prevent each, and why conversational AI interviews are structurally resistant to the junk that plagues panel surveys.
Survey Fatigue: Why It's Getting Worse (And How AI Interviews Solve It)
Survey fatigue is driving response rates to historic lows. This guide explains why it is happening, what it costs your research, and how AI-moderated interviews deliver better data without burning out respondents.
Survey Fraud & Respondent Quality: How to Detect Fake and Low-Effort Responses (2026)
Between 5% and 26% of survey responses are fraudulent, and AI-generated answers now pass standard quality checks. Learn the warning signs, the detection tactics that still work, and how Koji's conversational quality gate filters bad data before it reaches your report.
Synthetic Users in Research: Validity, Bias, and When AI Personas Are (and Aren't) Trustworthy
A research methodology guide to synthetic users — what they are, the documented bias problems (sycophancy, sign-flipping, shallow insights), the legitimate use cases, and why real AI-moderated interviews are now fast enough that the synthetic-vs-real tradeoff has fundamentally shifted.