{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-07T14:23:36.727Z"},"content":[{"type":"blog","id":"403c7ed5-d3ca-4302-911e-9fb40f292d3d","slug":"ai-generated-survey-responses-2026","title":"Are Your Survey Open-Ends Written by AI? The 2026 Data-Quality Crisis","url":"https://www.koji.so/blog/ai-generated-survey-responses-2026","summary":"Peer-reviewed 2025 evidence shows widespread LLM contamination of online survey data: 34 percent of participants self-report using LLMs for open-ends (Zhang, Xu and Alvero, Sociological Methods & Research), 45 percent showed copy-paste behavioural signatures (Rilla et al.), and an autonomous AI agent passed 99.8 percent of standard attention checks at about five cents per response (Westwood, PNAS). Detection is structurally losing because the cost asymmetry favours attackers and harder checks bias real samples. LLM-mediated text is uniquely dangerous because it adds homogeneity rather than noise, counterfeiting data saturation. Defences that work are structural: real-time modes, unpredictable adaptive follow-ups, episode-specific questions, verified own-customer recruitment, and retained auditable recordings.","content":"**Short answer:** yes, some of them almost certainly are. In a 2025 study published in *Sociological Methods & Research*, 34 percent of online research participants reported using large language models to help them answer open-ended survey questions. In a separate 2025 paper in *PNAS*, an autonomous AI agent built from a simple prompt passed **99.8 percent** of standard attention checks - and produced coherent, reasoned survey responses for about five cents each.\n\nThis is not the familiar problem of bots and click-farms, which the industry has been managing for a decade. It is a different problem with a different shape, and most of the defences currently sold as data-quality tooling were designed for the old one. This piece sets out what the published evidence shows, why the detection arms race is structurally unwinnable, and which research designs actually change the economics.\n\n## What the 2025 evidence actually says\n\n### 34 percent admit to it\n\nSimone Zhang, Janet Xu and AJ Alvero surveyed research participants recruited from a widely used social science subject platform and asked them directly. **Thirty-four percent reported using LLMs to help answer open-ended survey questions.** That is self-report, which means it is a floor, not a ceiling.\n\nThe more troubling half of their finding is what the assisted answers looked like. LLM-shaped responses were measurably **more homogeneous and more positive** than human-written ones, and the effect was strongest exactly where researchers care most: sensitive questions describing social groups. The authors warn that this homogenization \"may mask important underlying social variation in attitudes and beliefs.\"\n\n### Up to 45 percent showed the behavioural signature\n\nRaluca Rilla, Tobias Werner, Hiromu Yakura, Iyad Rahwan and Anne-Marie Nussberger, writing in late 2025, ran a pilot that logged behaviour rather than trusting self-report. On an open-ended question page, **45 percent of participants performed copying or pasting actions** consistent with routing the question through an external model.\n\nTheir most useful contribution is a taxonomy. They separate three distinct phenomena:\n\n- **Partial LLM mediation** - a real person uses a model for translation, phrasing or idea generation. A human is present; the words are not theirs.\n- **Full LLM delegation** - an agentic model completes the study with little or no human oversight. There is no respondent at all.\n- **LLM spillover** - real, unassisted humans behave differently because they now expect AI to be involved.\n\nThey also note that automated classifiers, behavioural heuristics and self-reports all produced **different prevalence estimates on the same data**, which tells you how unreliable any single detection method currently is. Earlier work by Veselovsky and colleagues found that even when LLM use was explicitly prohibited, up to 24 percent of crowdworkers used one anyway.\n\n### The attention check is dead\n\nThe most important paper is Sean Westwood's, published in *PNAS* in 2025 under the title \"The potential existential threat of large language models to online survey research.\" It has already been cited more than 60 times.\n\nWestwood built a deliberately minimal autonomous respondent and ran **43,800 distinct evaluations** across 139 questions and 6,700 trials with randomly generated demographic personas. The results:\n\n- **99.8 percent pass rate on attention checks** - 10 errors across 6,000 trials spanning 20 distinct check types drawn from prominent published papers.\n- **Perfect performance on 18 of the 20 check categories.**\n- **Zero errors** on impossible-event probes, the questions designed to catch respondents claiming implausible experiences.\n- **97.7 percent strategic refusal** on reverse shibboleth tasks - the agent concealed superhuman ability rather than revealing it, which is to say it knew what it was being tested for.\n- Instruction-following tasks, reading comprehension, logical reasoning probes, behavioural flags, response-pattern analysis, and audio and video attention checks all failed to catch it.\n\nTwo further findings should worry anyone running commercial research. The agent **correctly inferred the study hypothesis in 84 percent of trials** and shifted its answers accordingly, producing a 22.2 percentage-point increase in hypothesis-confirming responses. And Westwood calculated that between **10 and 52 synthetic responses** would have been enough to flip predictions in 2024 election polls.\n\nWestwood is explicit that this is \"a lower bound on the potential threat,\" because what he built was a proof of concept.\n\n### The evidence at a glance\n\n| Study | Method | Headline finding |\n|---|---|---|\n| Zhang, Xu and Alvero (2025), *Sociological Methods & Research* | Direct self-report from online research participants | 34% used LLMs on open-ended questions; answers more homogeneous and more positive |\n| Rilla, Werner, Yakura, Rahwan and Nussberger (2025) | Behavioural logging of copy and paste on an open-ended page | 45% showed LLM-consistent behaviour; classifiers, heuristics and self-report disagreed on prevalence |\n| Veselovsky and colleagues | Crowdworkers explicitly prohibited from using LLMs | Up to 24% used one anyway |\n| Westwood (2025), *PNAS* | 43,800 evaluations of an autonomous AI respondent across 139 questions | 99.8% attention-check pass rate; about $0.05 per response |\n\n## Why detection is the wrong battlefield\n\n### The economics are inverted\n\nWestwood puts the cost of a synthetic response at roughly **five cents** using commercial APIs, against a typical survey payout of $1.50 - a margin above **96.8 percent**. Run on locally hosted open-weight models, the marginal cost approaches zero.\n\nSurvey fraud used to be labour-limited. A human click-farmer could only complete so many surveys per hour, so quality controls that raised the cost per completion worked. That lever no longer exists. Any defence whose cost scales with the volume of responses is now on the losing side of an exponential.\n\n### Harder checks damage your real sample\n\nThe instinctive response - write trickier attention checks - fails twice. It fails against the machine, as Westwood demonstrated across 20 check types. And it succeeds against exactly the wrong people: increasingly complex comprehension and instruction-following tasks screen out respondents with lower literacy or reading speed, biasing your sample while catching nothing.\n\nWestwood argues against the arms race directly, noting that \"the vulnerability exists because current data-quality safeguards were designed for a different era.\" His recommendations are structural rather than technological: panel transparency requirements covering validation frequency, throttling, professionalism metrics and location verification; and a shift away from low-barrier convenience samples toward deeply vetted panels, address-based sampling and face-to-face contact. Our guide to [attention check questions](/docs/attention-check-questions) remains useful for catching genuinely inattentive humans - just do not mistake it for a defence against a competent model.\n\n### The part nobody flags as fraud\n\nFull delegation is fraud, and in principle a panel can be held responsible for it. Partial mediation is not. A real, verified, honest participant who pastes the question into a chatbot because English is their second language, or because they want to give you a thoughtful answer, has broken no rule you wrote down.\n\nThat respondent is real. Their identity checks out. Their location checks out. And the text in your dataset is not their voice. No fraud-detection product on the market will flag them, because nothing fraudulent happened. This is why framing the issue as \"survey fraud\" - the framing in most vendor material, including our own [survey fraud and respondent quality](/docs/survey-fraud-respondent-quality) guide - captures maybe half the problem.\n\n## The most dangerous property: it counterfeits saturation\n\nOrdinary bad data adds noise. Straight-lining, gibberish and random clicking add variance, and variance is visible - it widens intervals, it looks wrong, an analyst notices.\n\nLLM-mediated text does the opposite. It adds **agreement**. Zhang and colleagues found the assisted responses were more homogeneous, and homogeneity is precisely the pattern qualitative researchers have been trained to treat as a validation signal. When the twelfth interview surfaces nothing new, we call that [data saturation](/docs/data-saturation-qualitative-research) and conclude the sample was sufficient.\n\nA corpus polluted with model-generated open-ends reaches apparent saturation faster than a clean one, because the same model, prompted in similar ways, produces similar answers. **The contamination does not look like an error. It looks like proof that you were right.** A team can run a rigorous process, hit saturation on schedule, get sign-off, and ship - having been agreed with by a language model reflecting its own training data back at them.\n\nThis is also why \"we cross-checked the themes and they were consistent\" is no longer reassuring, and why the second-order effect on positivity matters commercially. Model-generated text skews positive. If a slice of your open-ends is model-written, your satisfaction verbatims are biased upward and your complaint themes are thinned - in the direction that makes a product team feel good.\n\n## What actually changes the economics\n\nNo design is immune. But some designs make the attack expensive instead of free.\n\n**1. Close the side channel with real time.** A typed text box permits an unbounded, unobserved pause. The respondent can alt-tab, paste, wait, copy and return, and nothing about the artifact records it. A live spoken conversation collapses that window: the moderator is waiting, the silence is audible, and the response arrives at the speed of speech. This does not make delegation impossible, but it removes the free, invisible version of it.\n\n**2. Make the questions unpredictable.** Westwood's agent worked so well partly because a questionnaire is a fixed, known artifact. Adaptive follow-ups generated from what the person just said cannot be pre-solved, because they do not exist until the answer does. An interview that asks \"you said the pricing page felt dishonest - which part,\" in response to an unscripted answer, is a much harder target than question 14 of 40. This is what Koji's [AI follow-up probing](/docs/ai-probing-guide) is doing structurally, whatever else it is doing for depth.\n\n**3. Ask for specifics only the account holder has.** Move from attitudes to episodes. \"What happened the last time you tried to export a report\" has a right answer that lives in your own logs. Generic questions invite generic answers, which is exactly what a model is best at producing.\n\n**4. Recruit from your own base, not an open panel.** Westwood's structural recommendation - deeply vetted panels over low-barrier convenience samples - maps cleanly onto commercial research. Your customers are identity-verified by your billing system, have a stake in the product, and are not being paid per completion. When you do buy sample, ask the questions in [how to choose a sample provider](/blog/how-to-choose-sample-provider-esomar-37-2026).\n\n**5. Keep the raw artifact, not just the text.** If your evidence is a stored recording and transcript tied to a session, a reviewer can check whether a human said it. If your evidence is a row in a spreadsheet, there is nothing left to audit. Our [survey data quality guide](/docs/survey-data-quality-guide) covers the operational checks; the point here is to retain something worth checking.\n\n**6. Report coverage, not just proportions.** For open-ended material, \"7 of 22 participants raised this unprompted\" is a more honest and more robust claim than a percentage, and it degrades gracefully if a few responses turn out to be contaminated.\n\n## Where this leaves the survey\n\nNone of this means surveys are finished. Closed-ended instruments remain the right tool for measurement, and the tradeoffs are laid out in [open-ended vs closed-ended questions](/docs/open-ended-vs-closed-ended-questions). What has changed is the status of the **free-text box**, which for thirty years was the cheap way to get a little qualitative colour attached to your quant. That box is now the least defensible element of the instrument, and it is the one most likely to be quoted in a deck.\n\nThe honest read of the 2025 literature is that unsupervised, low-barrier, text-only data collection has lost the presumption of authenticity. Westwood says the assumption that logically coherent responses come from humans \"is now untenable.\" That is a strong claim from a peer-reviewed venue, and it should change what teams are willing to bet on.\n\n## Frequently Asked Questions\n\n### How common are AI-generated survey responses in 2026?\n\nThe best available estimates come from 2025. Zhang, Xu and Alvero found 34 percent of online research participants self-reported using LLMs to answer open-ended questions. Rilla and colleagues logged copy-paste behaviour consistent with LLM use in 45 percent of participants on an open-ended page. Veselovsky and colleagues found up to 24 percent used models even when explicitly told not to. Estimates vary because detection methods disagree.\n\n### Can AI detection tools catch LLM-written survey answers?\n\nNot reliably. Rilla and colleagues found that automated classifiers, behavioural heuristics and self-reports produced different prevalence estimates on the same data. Commercial detectors also produce false positives against non-native speakers and formal writers, so using them to reject responses introduces its own sampling bias. Treat detection as a weak signal, not a gate.\n\n### Do attention checks still work?\n\nAgainst inattentive humans, yes. Against a competent model, no. Westwood reported a 99.8 percent pass rate for an autonomous AI respondent across 20 distinct attention-check types, with perfect performance on 18 of them and zero errors on impossible-event probes. Making checks harder mainly screens out real respondents with lower literacy.\n\n### Why is AI-generated text worse than ordinary bad survey data?\n\nBecause it adds agreement rather than noise. Model-generated answers are more homogeneous and more positive than human ones, so contaminated data reaches apparent saturation faster and looks more consistent. It counterfeits the exact pattern researchers treat as evidence that a finding is solid, which makes it far harder to notice than gibberish.\n\n### Are voice interviews safer than text surveys?\n\nThey are harder and more expensive to fake, not immune. A live spoken conversation removes the unobserved pause a text box allows, and unscripted follow-ups cannot be pre-solved because they are generated from what the participant just said. Combined with retained recordings and transcripts, that gives you an artifact a reviewer can actually audit.\n\n### What should we change first?\n\nStop treating free-text boxes as qualitative evidence, and move the questions that matter into a mode where a real person has to respond in real time to something they could not anticipate. Then recruit from your own verified customer base where possible, and keep the raw recording so any finding can be traced back to a human being.\n\n## Get evidence you can trace to a person\n\nKoji runs AI-moderated interviews by voice or text with real participants. Follow-ups are generated live from what each person actually says, so there is no fixed questionnaire to pre-solve. Every theme in the report resolves back to a named interview and a moment in a retained transcript, and structured questions - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` - give you the numbers alongside the reasoning, as described in the [structured questions guide](/docs/structured-questions-guide).\n\nYou can also read our evidence review on whether [people are candid with an AI interviewer](/blog/will-customers-open-up-to-ai-interviewer-2026), which turns out to be a separate question with a much older literature.\n\n**[Start free with 10 credits](https://www.koji.so/signup)** - enough to run a real study and see what a traceable transcript looks like.\n","category":"Research","lastModified":"2026-08-07T03:26:00.039597+00:00","metaTitle":"Are Your Survey Open-Ends Written by AI? The 2026 Data-Quality Crisis","metaDescription":"34% of online research participants use LLMs on open-ended questions, and an AI agent passed 99.8% of attention checks in a 2025 PNAS study. What the evidence shows, why detection is losing, and which designs still work.","keywords":["ai generated survey responses","chatgpt survey responses","llm survey data quality","fake survey responses","survey data integrity","attention checks ai","online panel data quality"],"aiSummary":"Peer-reviewed 2025 evidence shows widespread LLM contamination of online survey data: 34 percent of participants self-report using LLMs for open-ends (Zhang, Xu and Alvero, Sociological Methods & Research), 45 percent showed copy-paste behavioural signatures (Rilla et al.), and an autonomous AI agent passed 99.8 percent of standard attention checks at about five cents per response (Westwood, PNAS). Detection is structurally losing because the cost asymmetry favours attackers and harder checks bias real samples. LLM-mediated text is uniquely dangerous because it adds homogeneity rather than noise, counterfeiting data saturation. Defences that work are structural: real-time modes, unpredictable adaptive follow-ups, episode-specific questions, verified own-customer recruitment, and retained auditable recordings.","aiKeywords":["ai generated survey responses","llm survey pollution","attention checks","data saturation","survey data quality","westwood pnas 2025"],"aiContentType":"guide","faqItems":[{"answer":"The best available estimates come from 2025. Zhang, Xu and Alvero found 34 percent of online research participants self-reported using LLMs to answer open-ended questions. Rilla and colleagues logged copy-paste behaviour consistent with LLM use in 45 percent of participants on an open-ended page. Veselovsky and colleagues found up to 24 percent used models even when explicitly told not to. Estimates vary because detection methods disagree.","question":"How common are AI-generated survey responses in 2026?"},{"answer":"Not reliably. Rilla and colleagues found that automated classifiers, behavioural heuristics and self-reports produced different prevalence estimates on the same data. Commercial detectors also produce false positives against non-native speakers and formal writers, so using them to reject responses introduces its own sampling bias. Treat detection as a weak signal, not a gate.","question":"Can AI detection tools catch LLM-written survey answers?"},{"answer":"Against inattentive humans, yes. Against a competent model, no. Westwood reported a 99.8 percent pass rate for an autonomous AI respondent across 20 distinct attention-check types, with perfect performance on 18 of them and zero errors on impossible-event probes. Making checks harder mainly screens out real respondents with lower literacy.","question":"Do attention checks still work?"},{"answer":"Because it adds agreement rather than noise. Model-generated answers are more homogeneous and more positive than human ones, so contaminated data reaches apparent saturation faster and looks more consistent. It counterfeits the exact pattern researchers treat as evidence that a finding is solid, which makes it far harder to notice than gibberish.","question":"Why is AI-generated text worse than ordinary bad survey data?"},{"answer":"They are harder and more expensive to fake, not immune. A live spoken conversation removes the unobserved pause a text box allows, and unscripted follow-ups cannot be pre-solved because they are generated from what the participant just said. Combined with retained recordings and transcripts, that gives you an artifact a reviewer can actually audit.","question":"Are voice interviews safer than text surveys?"},{"answer":"Stop treating free-text boxes as qualitative evidence, and move the questions that matter into a mode where a real person has to respond in real time to something they could not anticipate. Then recruit from your own verified customer base where possible, and keep the raw recording so any finding can be traced back to a human being.","question":"What should we change first?"}],"relatedTopics":["survey data quality","ai generated responses","research integrity","online panels","attention checks","qualitative research"]}],"pagination":{"total":1,"returned":1,"offset":0}}