{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-02T09:36:00.056Z"},"content":[{"type":"documentation","id":"e6e3a84e-61e6-441f-977e-c2ef79faf6ff","slug":"ai-detector-false-positives-research","title":"Why You Cannot Gate Research on an AI Detector","url":"https://www.koji.so/docs/ai-detector-false-positives-research","summary":"AI text detectors cannot be used as an exclusion gate for research responses. Liang et al. (Patterns, 2023) evaluated seven detectors and found an average false positive rate of 61.22 percent on 91 TOEFL essays by non-native English writers versus 5.19 percent on US eighth-grade essays; 97.80 percent of TOEFL essays were flagged by at least one detector. The mechanism is text predictability, not authorship, which is why simplifying native essays raised misclassification to 56.65 percent and why a rewrite prompt cut false positives to 11.77 percent -- the tools are biased and bypassable through the same property. Worked arithmetic: a detector gate applied to 1,000 honest responses with 20 percent non-native writers discards 164 responses and cuts non-native representation to 9.3 percent. Use batch-level lexical diversity monitoring instead.","content":"Do not use an AI text detector to decide which research responses to keep. Published evaluations show these tools misclassify non-native English writing as AI-generated at rates far above their rates for native writing, while being trivially bypassable by anyone actually trying to evade them. A detector-based exclusion rule therefore removes honest participants, concentrates those removals among non-native speakers, and still fails to catch a determined user. It makes your sample worse in a direction you cannot see.\n\n## The short answer\n\nAn AI detector fails in both directions at the same time:\n\n- **False positives are high and unevenly distributed.** Honest writing gets flagged, and non-native English writing gets flagged far more often than native writing.\n- **False negatives are cheap to produce.** A single prompt instructing a model to write in a less predictable style defeats the detector.\n\nA tool that wrongly accuses the compliant while waving through the evasive is not a weak control. It is a control pointed the wrong way. Use detector output, if at all, as a soft prompt to look at an instrument, never as a rule that removes a person's data.\n\n## The core evidence\n\nThe standard reference is Liang, Yuksekgonul, Mao, Wu and Zou, \"GPT detectors are biased against non-native English writers,\" published in *Patterns* in 2023. The authors evaluated seven widely-used GPT detectors against writing samples from native and non-native English writers, using a set of ninety-one TOEFL essays written by non-native speakers and a comparison set of essays by US eighth-grade students.\n\nThe headline results:\n\n| Sample | Detector behaviour |\n| --- | --- |\n| 91 TOEFL essays (non-native) | Average false positive rate 61.22 percent |\n| Same essays, unanimous verdicts | All seven detectors flagged 18 of the 91 essays, 19.78 percent, as AI-authored |\n| Same essays, any single detector | 89 of the 91 essays, 97.80 percent, flagged by at least one detector |\n| US eighth-grade essays (native) | Average of 5.19 percent across detectors |\n\nThe authors summarise the central finding plainly in the abstract: the detectors \"consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified.\"\n\nTwo further results sharpen the point.\n\n**The bias tracks linguistic simplicity, not authorship.** When the native-speaker essays were simplified to resemble non-native writing, the misclassification rate rose to 56.65 percent. So the detectors are not detecting machine authorship. They are detecting constrained vocabulary and predictable sentence construction, and then reporting that as machine authorship. The authors conclude that GPT detectors \"may unintentionally penalize writers with constrained linguistic expressions.\"\n\n**The same mechanism makes them bypassable.** Prompting a model to rewrite its output in richer language cut the average false positive rate on the TOEFL essays from 61.22 percent to 11.77 percent, a decrease of 49.45 percent. The identical lever that rescues an honest non-native writer also hides a genuinely AI-written response. You cannot tune your way out of this, because both failures are the same measurement.\n\n## The arithmetic that settles it\n\nPrevalence rates and percentages are abstract. Run them through a sample and the problem becomes concrete.\n\nTake 1,000 open-ended responses, of which 200 come from non-native English writers and 800 from native writers. To isolate the false-positive behaviour, assume for the moment that nobody used an LLM at all -- every response is the participant's own work. Apply a rule that drops anything a detector flags, using the measured rates above:\n\n- Non-native writers flagged: 200 x 0.6122 = 122\n- Native writers flagged: 800 x 0.0519 = 42\n- Total honest responses discarded: 164\n\nNow look at who is left. Your surviving sample is 78 non-native writers and 758 native writers, 836 responses in total. Non-native representation has fallen from 200 of 1,000, which is 20.0 percent, to 78 of 836, which is 9.3 percent. You have more than halved the share of non-native voices in your findings.\n\nAnd 122 of the 164 responses you discarded, which is 74.4 percent of your exclusions, came from a group that made up only 20 percent of the sample.\n\nThat is a [sampling bias](/docs/sampling-bias-research) you created yourself, with a filter you probably did not document, acting hardest on participants whose perspective you were least likely to have enough of already. If your product has international users, this rule quietly deletes them from your evidence base. Note too that the real situation is worse than this calculation, because some genuine LLM users evade the filter entirely -- so you pay the full cost in lost honest responses without getting the benefit you bought it for.\n\n## Why the asymmetry is unfixable by threshold\n\nIt is tempting to assume a stricter threshold solves this. It does not, for a structural reason.\n\nThe signal these tools lean on is text predictability -- roughly, how unsurprising each word is given the words before it. Language models emit low-surprise text because that is what they are optimised to do. But a human writing in a second language also produces lower-surprise text, for an entirely unrelated reason: a smaller active vocabulary and more conventional constructions. The same is true of people writing in a hurry, people writing on a phone, people with limited formal education in the language, and people writing in a plain professional register because they think that is what you want.\n\nA threshold can trade false positives against false negatives, but it cannot separate two populations that genuinely overlap on the measured quantity. Raise the threshold and you lose the ability to detect anything; lower it and you sweep up more honest non-native writers. There is no setting at which the tool becomes a fair gate, because the construct it measures is not the construct you care about.\n\nThis is a textbook [screener accuracy](/docs/screener-accuracy-positive-predictive-value) problem, and the same base-rate logic applies: when the flagged population is dominated by honest writers, a positive result tells you very little about the individual it was applied to.\n\n## What detector output is legitimately good for\n\nNot nothing, as long as it never touches an individual's data:\n\n- **As a trend on your own instrument.** If the flag rate on one question jumps between waves while your population has not changed, that question has probably started inviting outsourcing. Investigate the question.\n- **As a reason to improve question design.** A high flag rate is weak evidence that your open-ends are generic enough to be answerable without the participant's own experience. That is actionable and blames nobody.\n- **Never as an exclusion criterion, a payment decision, or an accusation.** Withholding compensation on a detector score is both unjust and, given a 61.22 percent false positive rate on non-native writing, very likely to be wrong.\n\nThe defensible alternative is to measure the batch rather than the person: track lexical diversity and near-duplicate rates across a wave and act on movement in the aggregate. That diagnostic carries no individual accusation and no disparate impact. Because Koji maps every answer to a stable question ID from the research brief through to report aggregation, each wave measures the same question, which is what makes the comparison trustworthy.\n\n## How Koji handles this\n\nKoji's position is that provenance is won at collection time, not recovered afterwards by a classifier.\n\n- **Adaptive probing tests knowledge, not style.** Koji's AI interviewer generates each follow-up from the previous answer, so a response is validated by whether the participant can supply specifics and stay consistent -- not by how predictable their prose is. That test is fair to a non-native speaker, who may write plainly and still know exactly what happened on Tuesday.\n- **Voice mode sidesteps the question.** Koji runs interviews in voice as well as text. A spoken answer has no paste surface, so there is nothing for a detector to adjudicate.\n- **Structured questions carry the countable load.** With six question types available -- `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` -- the ratings and selections live in typed fields. That shrinks how much your open-ends have to do, and a shorter open-end is both less tempting to outsource and less likely to be flagged for plain phrasing.\n- **Quality scoring rates substance, not fluency.** Koji scores each interview 1-5 across relevance, depth and coverage. An articulate conversation that never reaches a specific is marked down on depth. A blunt, plainly-worded conversation full of concrete detail scores well. This is deliberately the inverse of what a detector rewards.\n- **Exclusions stay visible.** Because answers are mapped to stable question IDs from the brief through to report aggregation, whatever filtering you do apply is explicit and countable rather than an invisible edit to your sample.\n\n## Common mistakes\n\n- **Using a detector score as a payment or exclusion rule.** This is the error the Liang results most directly forbid. It is unfair to individuals and biases your sample.\n- **Assuming a confident percentage is a reliable one.** These tools return decisive-looking scores. A 99 percent confidence display does not correspond to a 1 percent error rate.\n- **Testing the tool only on native-speaker writing.** A detector will look excellent at a 5.19 percent false positive rate and then behave completely differently on your actual international sample.\n- **Forgetting the bypass.** Any participant who wants to evade detection can, with one extra prompt. Your gate is therefore selecting against honesty. Koji treats provenance as something won at collection time through adaptive probing and voice mode, rather than guessed at afterwards from prose style.\n- **Not documenting the filter.** If you do exclude responses, record the rule and report how many it removed, exactly as you would for any other [response bias](/docs/survey-response-bias) adjustment. An undocumented filter is an unreproducible study.\n\n## Frequently asked questions\n\n### How accurate are AI content detectors on survey responses?\n\nFar less accurate than their interfaces suggest, and their errors are not evenly spread. In the Liang evaluation, seven detectors produced an average false positive rate of 61.22 percent on ninety-one TOEFL essays written by non-native English speakers, against an average of 5.19 percent on essays by US eighth-grade students. Accuracy that depends this heavily on who wrote the text is not usable as a gate.\n\n### Why do AI detectors flag non-native English writers so often?\n\nBecause they measure text predictability rather than authorship. Language models produce low-surprise text by design, and so does a competent writer working in a second language with a smaller active vocabulary and more conventional sentence structure. The two populations overlap on the exact quantity being measured. The Liang study demonstrated this directly: simplifying native-speaker essays to resemble non-native writing pushed misclassification up to 56.65 percent.\n\n### Can I just raise the detection threshold to be safe?\n\nNo. A threshold trades one error type for the other, but it cannot separate groups that genuinely overlap on the measured signal. Raising it blinds the tool; lowering it catches more honest writers. There is no value at which the tool becomes a fair exclusion rule, which is why the recommendation is to stop using it as one rather than to tune it.\n\n### Is there any defensible way to detect AI-written responses?\n\nNot at the level of an individual response, no. What is defensible is measuring a whole batch: track lexical diversity and near-duplicate rates wave over wave and act on movement in the aggregate. That approach accuses nobody, has no disparate impact on non-native speakers, and still tells you when a question has started attracting generic answers.\n\n### Should I tell participants I am running their answers through a detector?\n\nIf you are going to act on the output in any way that affects them, yes, and that disclosure requirement is itself a good reason not to. A far better use of the same honesty is to explain up front that you want their own words and why, then ask directly and without penalty whether they used an AI tool. That gives you a usable covariate instead of a contested score.\n\n### Does this mean AI-assisted responses are not a real problem?\n\nThey are a real problem. The evidence on participant LLM use and on the resulting homogenization of open-ended text is solid. The argument here is narrower and only about the remedy: detection at the individual level is the wrong instrument for it. Fix it with question design that demands participant-specific detail, voice collection where risk is high, and corpus-level monitoring.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - move quantitative load into typed questions so open-ends stay short\n- [Screener Accuracy and Positive Predictive Value](/docs/screener-accuracy-positive-predictive-value) - the same base-rate logic applied to recruitment\n- [Sampling Bias in Research](/docs/sampling-bias-research) - what an undocumented exclusion filter does to your findings\n- [Survey Response Bias](/docs/survey-response-bias) - documenting and reporting adjustments\n- [Synthetic Users in Research](/docs/synthetic-users-research-methodology) - the researcher-side version of the AI question\n- [Survey Fraud and Respondent Quality](/docs/survey-fraud-respondent-quality) - controls that do work, against a different problem","category":"Research Methods","lastModified":"2026-10-02T03:47:16.127357+00:00","metaTitle":"Why You Cannot Gate Research on an AI Detector (2026)","metaDescription":"AI detectors flagged 61.22% of non-native English essays as AI-written. Why detector-based exclusion biases your sample.","keywords":["ai detector false positives","ai detection accuracy","gpt detector bias","non-native english writers","ai content detector research","detector exclusion bias"],"aiSummary":"AI text detectors cannot be used as an exclusion gate for research responses. Liang et al. (Patterns, 2023) evaluated seven detectors and found an average false positive rate of 61.22 percent on 91 TOEFL essays by non-native English writers versus 5.19 percent on US eighth-grade essays; 97.80 percent of TOEFL essays were flagged by at least one detector. The mechanism is text predictability, not authorship, which is why simplifying native essays raised misclassification to 56.65 percent and why a rewrite prompt cut false positives to 11.77 percent -- the tools are biased and bypassable through the same property. Worked arithmetic: a detector gate applied to 1,000 honest responses with 20 percent non-native writers discards 164 responses and cuts non-native representation to 9.3 percent. Use batch-level lexical diversity monitoring instead.","aiPrerequisites":["Familiarity with open-ended survey questions","Basic understanding of false positive and false negative rates"],"aiLearningOutcomes":["State the measured false-positive rates for GPT detectors on non-native versus native writing","Explain why text predictability is not a measure of authorship","Compute the sample-composition damage a detector gate causes","Identify the narrow legitimate uses of detector output","Choose corpus-level diagnostics over individual-level detection"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}