{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-02T09:40:28.138Z"},"content":[{"type":"documentation","id":"de4ca0ae-8481-4141-bb1a-8720deaff4bb","slug":"ai-assisted-open-ended-responses","title":"AI-Assisted Answers: What to Do When Participants Let an LLM Write Their Open-Ends","url":"https://www.koji.so/docs/ai-assisted-open-ended-responses","summary":"AI-assisted open-ended answers are a distinct failure class from fraud, satisficing and synthetic users: the participant is real and eligible but the prose is not theirs, so eligibility screeners and attention checks all pass. Zhang, Xu and Alvero (Sociological Methods & Research, 2025) found 34 percent of participants on a paid research-subject platform reported using LLMs for open-ends, and their simulations found LLM responses more homogeneous and positive. The expensive cost is variance collapse, which makes data look better while destroying the tail signal open-ends exist to capture. Correct response is question design anchored in participant-specific episodes, corpus-level diversity monitoring, and a non-punitive disclosure question -- never detector-based exclusion.","content":"When a participant pastes your open-ended question into a chatbot and pastes the answer back, you have a validity problem, not a fraud problem. The person is real, eligible and willing. Only the words are not theirs. That distinction decides your entire response: detection and exclusion is the wrong instinct, and question design plus corpus-level monitoring is the right one.\n\n## The short answer\n\nTreat AI-assisted open-ends as a measurement failure you caused, not a betrayal the participant committed. Three moves, in order of value:\n\n1. **Design questions a generic answer cannot answer.** Ask for a specific episode with a date, a sequence, or a number the participant alone could know. A language model can produce a fluent paragraph about onboarding friction; it cannot tell you what happened the Tuesday your billing page timed out.\n2. **Monitor the batch, not the person.** You cannot responsibly judge any single response, but you can measure whether a whole wave of open-ends has lost its variance. That is a corpus statistic, and it is defensible.\n3. **Stop rewarding length.** Most outsourcing is a rational response to a question that demanded an essay for a flat incentive.\n\nWhat you should not do is run responses through an AI detector and drop the flagged ones. The false-positive behaviour of those tools makes that indefensible, and it punishes exactly the participants you can least afford to lose.\n\n## This is not fraud, and the difference is load-bearing\n\nResearch operations already has vocabulary for bad responses, and AI-assisted answers fit none of the existing boxes:\n\n| Failure | Who is answering | What is wrong | Existing playbook |\n| --- | --- | --- | --- |\n| Survey fraud | A bot, or a person misrepresenting eligibility | The respondent is not who they claim | Identity and eligibility checks |\n| Satisficing or inattention | A real, eligible person | Effort is too low to be informative | Attention checks, shorter instruments |\n| Synthetic users | Nobody. The researcher substituted a model for people | There is no respondent at all | A methodology decision, made by you |\n| AI-assisted open-ends | A real, eligible, willing person | The person is genuine; the prose is not theirs | Mostly missing |\n\nThe practical consequence: every control built for the first three rows misses the fourth. An eligibility screener passes this participant, because they are eligible. An attention check passes them, because they are paying attention -- arguably more attention than someone typing six words. A fraud scorecard that looks for implausible speed can actively mislead you, because a participant who stops to prompt a chatbot may take longer, not less time, than an honest fast typist.\n\nSo the gap is real, and it is not covered by [survey fraud detection](/docs/survey-fraud-respondent-quality), [attention checks](/docs/attention-check-questions), or the [synthetic users](/docs/synthetic-users-research-methodology) debate, all of which answer different questions.\n\n## What the evidence actually shows\n\nThe best current measurement comes from Zhang, Xu and Alvero, published in *Sociological Methods & Research* in 2025 under the title \"Generative AI Meets Open-Ended Survey Responses: Research Participant Use of AI and Homogenization.\"\n\nTwo findings matter, and they come from two different parts of the study, which is worth keeping straight.\n\n**The prevalence figure.** In an original survey of research participants recruited from a popular online platform for sourcing social science research subjects, 34 percent reported using LLMs to help them answer open-ended survey questions.\n\nRead that population description carefully before you reuse the number. These are paid participants on a research-subject platform -- people for whom answering surveys is a repeated, compensated activity. That is the highest-exposure population there is. It is not a measurement of your own customers, and quoting \"34 percent of customers use AI\" would be a straightforward misreading. Use it as evidence that the behaviour is common and normalised in paid panels, and as a reason to design defensively everywhere else.\n\n**The homogenization finding.** This did not come from watching those participants. The authors ran simulations comparing human-written responses from three pre-ChatGPT studies against LLM-generated text, and found that LLM responses are more homogeneous and positive, particularly when they describe social groups in sensitive questions. Their stated concern is that these patterns may mask important underlying social variation in attitudes and beliefs among human subjects.\n\nThat is the sentence to tape to your monitor, because it names the real cost.\n\n## Why homogenization is the expensive part\n\nMost teams worry about AI-assisted answers being *fake*. The more serious problem is that they are *average*.\n\nQualitative research does not exist to find the modal opinion. If you wanted the modal opinion you would run a [structured scale question](/docs/structured-questions-guide) and count. Open-ends exist to surface the thing you did not think to ask about: the unusual workaround, the specific trigger, the minority use case that turns out to be a segment. That signal lives in the tail of the distribution.\n\nA language model is, by construction, a machine for producing the centre of a distribution. It is trained and tuned to emit the fluent, agreeable, broadly representative answer. So when a share of your open-ends come from a model, you do not get noise scattered around the truth -- you get mass pulled toward the middle, and the tail thins out.\n\nThis has an unpleasant property: it makes your data look *better*. Variance drops. Responses get longer and more articulate. Themes consolidate cleanly, and inter-coder agreement may even improve, because homogeneous text is easier to code consistently. Every surface signal of quality improves while the thing you were buying -- variation -- quietly disappears. Compare this with low-effort responses, which announce themselves as junk. AI-assisted responses are camouflaged as excellence.\n\nIf you take one operational lesson from this page: a wave of open-ends that suddenly reads better than usual deserves more scrutiny, not less.\n\n## The four conditions that invite outsourcing\n\nParticipants are not adversaries. They outsource when the instrument makes outsourcing the sensible move. Four design choices do most of the damage:\n\n### A question that rewards volume over specificity\n\n\"Tell us about your experience with our onboarding process\" is an essay prompt. It has no verifiable content, no anchor in time, and no way to be wrong. It is the single most outsourceable question shape there is. Contrast with \"What was the last thing that made you stop and ask someone for help?\" -- answerable in one sentence, and not answerable at all by a model that has never used your product.\n\n### A length cue that implies effort\n\nMinimum character counts, a large empty textarea, and progress copy like \"the more detail the better\" all signal that length is the currency. Participants who want to comply honestly, and who are tired, will reach for the fastest tool that produces length.\n\n### Flat compensation for variable effort\n\nIf every open-end pays the same regardless of difficulty, the rational participant minimises cost per unit of acceptable output. This is ordinary [satisficing](/docs/acquiescence-bias) economics operating on a new substrate. The chatbot did not create the incentive; it removed the last friction from acting on it.\n\n### An instrument long enough to exhaust goodwill\n\nOutsourcing concentrates late in long instruments, for the same reason [survey fatigue](/docs/survey-fatigue) and breakoff do. A participant who answered your first four open-ends honestly and is staring at a fifth has a strong motive to finish by other means. The fix is the one you already know: ask for less. See [ideal survey length](/docs/ideal-survey-length-guide) for the trade-offs.\n\n## What to do instead\n\n### Anchor every open-end in something only this person knows\n\nThe strongest defence is epistemic, not technical. Ask for episodes, not opinions:\n\n- \"Walk me through the last time you tried to do X. What happened first?\"\n- \"What did you do immediately after that?\"\n- \"Who else was involved, and what did they say?\"\n- \"What did you expect to happen instead?\"\n\nA model can invent a plausible answer to each of these, but it cannot keep a specific invented episode internally consistent across three unscripted follow-ups that depend on the previous answer. That is the mechanism that actually works, and it is why conversational formats hold up better than forms here.\n\n### Ask about AI use directly, and non-punitively\n\nA plain question near the end -- \"Did you use any AI tool to help write your answers today?\" with a neutral frame and no penalty attached -- costs you one question and gives you a covariate you can analyse with. It will under-report. It is still far better than a detector score, because it is honest about its own uncertainty and it does not misclassify anyone. Pair it with clear framing up front about why you want their own words. Our guidance on [consent and disclosure wording](/docs/consent-form-wording-disclosure-research) applies directly: the wording is part of the instrument.\n\n### Measure the corpus\n\nTrack, wave over wave, the lexical diversity and near-duplicate rate of your open-ends. Falling diversity across a batch is a signal you can act on without accusing any individual. This is the defensible counterpart to detection, and it deserves its own treatment.\n\n### Never gate on a detector\n\nCovered at length separately, but the short version: the published false-positive rates on non-native English writing are high enough that a detector-based exclusion rule systematically removes non-native speakers from your sample. That is a sampling bias you introduced, and it is worse than the problem you were trying to fix.\n\n## How Koji handles this\n\nKoji's architecture attacks the cause rather than the symptom, because an AI-moderated conversation is a fundamentally harder target for outsourcing than a form.\n\n- **Adaptive follow-ups break invented episodes.** Koji's AI interviewer reads the previous answer and asks the next question from it. A pasted paragraph survives one exchange; it rarely survives being asked what happened next, who else was involved, and what the participant expected instead. The built-in `mom_test` framework encodes exactly this, with question patterns like \"Walk me through how you currently handle [task]\" and \"What happened after that?\" -- mechanism probes rather than opinion requests.\n- **Voice mode removes the copy-paste surface entirely.** In a spoken interview there is no textarea to paste into and no pause long enough to prompt a chatbot without it being obvious. Koji runs both voice and text, so you can choose the mode that matches your risk.\n- **Structured questions carry the quantitative load, so open-ends can be short.** Koji supports six question types -- `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no`. Move the ratings and the choices into typed questions and the open-ends no longer have to carry everything, which removes the length pressure that drives outsourcing in the first place.\n- **Per-interview quality scoring gives you an auditable signal.** Koji scores each interview 1-5 with a breakdown across relevance, depth and coverage. A conversation that is fluent but never lands on specifics scores low on depth regardless of word count, which is precisely the failure mode generic prose produces.\n- **Stable question IDs make the corpus comparable.** Because answers map back to specific question IDs from the brief through to report aggregation, you can compare the same question across waves and see diversity move.\n\n## Common mistakes\n\n- **Treating it as misconduct.** It drives the behaviour underground and poisons your disclosure question. The participant thinks they are being helpful.\n- **Reusing the 34 percent figure for your own customers.** It was measured on a paid research-subject platform. Your B2B customers are a different population with different incentives.\n- **Reading improved writing quality as improved data quality.** This is the trap. Fluency went up; variance went down. Only one of those is what you wanted.\n- **Adding a longer minimum character count.** This increases outsourcing pressure. It is the exact wrong lever.\n- **Excluding suspected responses silently.** If you drop responses on suspicion and do not record the rule, your sample is now shaped by an undocumented filter. At minimum, log the rule and report how many responses it removed, as you would for any other exclusion.\n\n## Frequently asked questions\n\n### Is a participant using AI to answer a survey question cheating?\n\nGenerally no, and framing it that way will cost you. The participant is real and eligible, and most people reach for a chatbot because the question asked for more writing than it was worth, not to deceive you. Treat it as a signal that your instrument created the incentive, and fix the question rather than policing the person.\n\n### How can I tell whether a specific response was written by an LLM?\n\nReliably, you cannot, and you should stop trying to decide this at the level of an individual response. There is no test with a false-positive rate low enough to justify acting against one participant. What you can do is measure a whole batch for falling lexical diversity and rising near-duplication, and anchor your questions in specifics so that a generic answer fails to be useful whether or not you can prove its origin.\n\n### Should I just ban AI use in my research?\n\nA stated expectation is worth having, because many participants will honour it once they understand you want their own words and why. An enforceable ban is not available to you, so do not build your data-quality plan on one. Combine a clear, non-punitive request with question design that makes generic answers worthless.\n\n### Does this affect voice interviews too?\n\nMuch less, and that is one of the strongest arguments for voice. A spoken conversation offers no paste target and no unobserved pause in which to prompt a model, and adaptive follow-ups force specifics in real time. Koji runs interviews in both voice and text, so a study with high outsourcing risk can simply be run in voice.\n\n### What should I do with responses I already suspect are AI-written?\n\nDo not delete them on suspicion. Keep them, code them, and compare them as a group against the rest of the batch on the dimensions you care about. If they are pulling your themes toward the centre you will see it as reduced variance, which is a finding you can report honestly. Any exclusion rule you do apply should be written down and its effect on sample size reported.\n\n### Does a longer, better-written answer mean better data?\n\nNo, and assuming so is the most expensive error on this page. Length and fluency are not quality. An answer that is articulate, balanced and generic tells you less than a short, awkward sentence about a specific Tuesday. Judge open-ends by whether they contain something only that person could have said.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and how to move quantitative load off your open-ends\n- [Survey Fraud and Respondent Quality](/docs/survey-fraud-respondent-quality) - the adjacent problem of fake and ineligible respondents\n- [Attention Check Questions](/docs/attention-check-questions) - catching low-effort responses without annoying real participants\n- [Synthetic Users in Research](/docs/synthetic-users-research-methodology) - when the researcher, not the participant, substitutes AI for people\n- [Survey Data Quality](/docs/survey-data-quality-guide) - the broader detection and prevention workflow\n- [Consent Form Wording](/docs/consent-form-wording-disclosure-research) - why your framing is part of the instrument","category":"Research Operations","lastModified":"2026-10-02T03:47:16.127357+00:00","metaTitle":"AI-Assisted Survey Answers: Handling LLM-Written Open-Ends (2026)","metaDescription":"Participants using ChatGPT for open-ends is a validity problem, not fraud. What the evidence shows and how to design around it.","keywords":["ai-generated survey responses","ai-assisted open-ends","llm survey answers","chatgpt survey responses","open-ended response quality","response homogenization"],"aiSummary":"AI-assisted open-ended answers are a distinct failure class from fraud, satisficing and synthetic users: the participant is real and eligible but the prose is not theirs, so eligibility screeners and attention checks all pass. Zhang, Xu and Alvero (Sociological Methods & Research, 2025) found 34 percent of participants on a paid research-subject platform reported using LLMs for open-ends, and their simulations found LLM responses more homogeneous and positive. The expensive cost is variance collapse, which makes data look better while destroying the tail signal open-ends exist to capture. Correct response is question design anchored in participant-specific episodes, corpus-level diversity monitoring, and a non-punitive disclosure question -- never detector-based exclusion.","aiPrerequisites":["Basic familiarity with open-ended survey questions","Understanding of response quality concepts"],"aiLearningOutcomes":["Distinguish AI-assisted responses from fraud, satisficing and synthetic users","Interpret the prevalence and homogenization evidence correctly, including its population limits","Identify the four instrument design choices that invite outsourcing","Write open-ended questions that a generic answer cannot satisfy","Choose corpus-level monitoring over individual-level detection"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}