Back to docs
Research Methods

Why You Cannot Gate Research on an AI Detector

AI text detectors misclassify non-native English writing at high rates while being trivially bypassable. Using one as an exclusion gate removes honest participants and biases your sample. Here is the evidence and the arithmetic.

Do not use an AI text detector to decide which research responses to keep. Published evaluations show these tools misclassify non-native English writing as AI-generated at rates far above their rates for native writing, while being trivially bypassable by anyone actually trying to evade them. A detector-based exclusion rule therefore removes honest participants, concentrates those removals among non-native speakers, and still fails to catch a determined user. It makes your sample worse in a direction you cannot see.

The short answer

An AI detector fails in both directions at the same time:

  • False positives are high and unevenly distributed. Honest writing gets flagged, and non-native English writing gets flagged far more often than native writing.
  • False negatives are cheap to produce. A single prompt instructing a model to write in a less predictable style defeats the detector.

A tool that wrongly accuses the compliant while waving through the evasive is not a weak control. It is a control pointed the wrong way. Use detector output, if at all, as a soft prompt to look at an instrument, never as a rule that removes a person's data.

The core evidence

The standard reference is Liang, Yuksekgonul, Mao, Wu and Zou, "GPT detectors are biased against non-native English writers," published in Patterns in 2023. The authors evaluated seven widely-used GPT detectors against writing samples from native and non-native English writers, using a set of ninety-one TOEFL essays written by non-native speakers and a comparison set of essays by US eighth-grade students.

The headline results:

SampleDetector behaviour
91 TOEFL essays (non-native)Average false positive rate 61.22 percent
Same essays, unanimous verdictsAll seven detectors flagged 18 of the 91 essays, 19.78 percent, as AI-authored
Same essays, any single detector89 of the 91 essays, 97.80 percent, flagged by at least one detector
US eighth-grade essays (native)Average of 5.19 percent across detectors

The authors summarise the central finding plainly in the abstract: the detectors "consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified."

Two further results sharpen the point.

The bias tracks linguistic simplicity, not authorship. When the native-speaker essays were simplified to resemble non-native writing, the misclassification rate rose to 56.65 percent. So the detectors are not detecting machine authorship. They are detecting constrained vocabulary and predictable sentence construction, and then reporting that as machine authorship. The authors conclude that GPT detectors "may unintentionally penalize writers with constrained linguistic expressions."

The same mechanism makes them bypassable. Prompting a model to rewrite its output in richer language cut the average false positive rate on the TOEFL essays from 61.22 percent to 11.77 percent, a decrease of 49.45 percent. The identical lever that rescues an honest non-native writer also hides a genuinely AI-written response. You cannot tune your way out of this, because both failures are the same measurement.

The arithmetic that settles it

Prevalence rates and percentages are abstract. Run them through a sample and the problem becomes concrete.

Take 1,000 open-ended responses, of which 200 come from non-native English writers and 800 from native writers. To isolate the false-positive behaviour, assume for the moment that nobody used an LLM at all -- every response is the participant's own work. Apply a rule that drops anything a detector flags, using the measured rates above:

  • Non-native writers flagged: 200 x 0.6122 = 122
  • Native writers flagged: 800 x 0.0519 = 42
  • Total honest responses discarded: 164

Now look at who is left. Your surviving sample is 78 non-native writers and 758 native writers, 836 responses in total. Non-native representation has fallen from 200 of 1,000, which is 20.0 percent, to 78 of 836, which is 9.3 percent. You have more than halved the share of non-native voices in your findings.

And 122 of the 164 responses you discarded, which is 74.4 percent of your exclusions, came from a group that made up only 20 percent of the sample.

That is a sampling bias you created yourself, with a filter you probably did not document, acting hardest on participants whose perspective you were least likely to have enough of already. If your product has international users, this rule quietly deletes them from your evidence base. Note too that the real situation is worse than this calculation, because some genuine LLM users evade the filter entirely -- so you pay the full cost in lost honest responses without getting the benefit you bought it for.

Why the asymmetry is unfixable by threshold

It is tempting to assume a stricter threshold solves this. It does not, for a structural reason.

The signal these tools lean on is text predictability -- roughly, how unsurprising each word is given the words before it. Language models emit low-surprise text because that is what they are optimised to do. But a human writing in a second language also produces lower-surprise text, for an entirely unrelated reason: a smaller active vocabulary and more conventional constructions. The same is true of people writing in a hurry, people writing on a phone, people with limited formal education in the language, and people writing in a plain professional register because they think that is what you want.

A threshold can trade false positives against false negatives, but it cannot separate two populations that genuinely overlap on the measured quantity. Raise the threshold and you lose the ability to detect anything; lower it and you sweep up more honest non-native writers. There is no setting at which the tool becomes a fair gate, because the construct it measures is not the construct you care about.

This is a textbook screener accuracy problem, and the same base-rate logic applies: when the flagged population is dominated by honest writers, a positive result tells you very little about the individual it was applied to.

What detector output is legitimately good for

Not nothing, as long as it never touches an individual's data:

  • As a trend on your own instrument. If the flag rate on one question jumps between waves while your population has not changed, that question has probably started inviting outsourcing. Investigate the question.
  • As a reason to improve question design. A high flag rate is weak evidence that your open-ends are generic enough to be answerable without the participant's own experience. That is actionable and blames nobody.
  • Never as an exclusion criterion, a payment decision, or an accusation. Withholding compensation on a detector score is both unjust and, given a 61.22 percent false positive rate on non-native writing, very likely to be wrong.

The defensible alternative is to measure the batch rather than the person: track lexical diversity and near-duplicate rates across a wave and act on movement in the aggregate. That diagnostic carries no individual accusation and no disparate impact. Because Koji maps every answer to a stable question ID from the research brief through to report aggregation, each wave measures the same question, which is what makes the comparison trustworthy.

How Koji handles this

Koji's position is that provenance is won at collection time, not recovered afterwards by a classifier.

  • Adaptive probing tests knowledge, not style. Koji's AI interviewer generates each follow-up from the previous answer, so a response is validated by whether the participant can supply specifics and stay consistent -- not by how predictable their prose is. That test is fair to a non-native speaker, who may write plainly and still know exactly what happened on Tuesday.
  • Voice mode sidesteps the question. Koji runs interviews in voice as well as text. A spoken answer has no paste surface, so there is nothing for a detector to adjudicate.
  • Structured questions carry the countable load. With six question types available -- open_ended, scale, single_choice, multiple_choice, ranking and yes_no -- the ratings and selections live in typed fields. That shrinks how much your open-ends have to do, and a shorter open-end is both less tempting to outsource and less likely to be flagged for plain phrasing.
  • Quality scoring rates substance, not fluency. Koji scores each interview 1-5 across relevance, depth and coverage. An articulate conversation that never reaches a specific is marked down on depth. A blunt, plainly-worded conversation full of concrete detail scores well. This is deliberately the inverse of what a detector rewards.
  • Exclusions stay visible. Because answers are mapped to stable question IDs from the brief through to report aggregation, whatever filtering you do apply is explicit and countable rather than an invisible edit to your sample.

Common mistakes

  • Using a detector score as a payment or exclusion rule. This is the error the Liang results most directly forbid. It is unfair to individuals and biases your sample.
  • Assuming a confident percentage is a reliable one. These tools return decisive-looking scores. A 99 percent confidence display does not correspond to a 1 percent error rate.
  • Testing the tool only on native-speaker writing. A detector will look excellent at a 5.19 percent false positive rate and then behave completely differently on your actual international sample.
  • Forgetting the bypass. Any participant who wants to evade detection can, with one extra prompt. Your gate is therefore selecting against honesty. Koji treats provenance as something won at collection time through adaptive probing and voice mode, rather than guessed at afterwards from prose style.
  • Not documenting the filter. If you do exclude responses, record the rule and report how many it removed, exactly as you would for any other response bias adjustment. An undocumented filter is an unreproducible study.

Frequently asked questions

How accurate are AI content detectors on survey responses?

Far less accurate than their interfaces suggest, and their errors are not evenly spread. In the Liang evaluation, seven detectors produced an average false positive rate of 61.22 percent on ninety-one TOEFL essays written by non-native English speakers, against an average of 5.19 percent on essays by US eighth-grade students. Accuracy that depends this heavily on who wrote the text is not usable as a gate.

Why do AI detectors flag non-native English writers so often?

Because they measure text predictability rather than authorship. Language models produce low-surprise text by design, and so does a competent writer working in a second language with a smaller active vocabulary and more conventional sentence structure. The two populations overlap on the exact quantity being measured. The Liang study demonstrated this directly: simplifying native-speaker essays to resemble non-native writing pushed misclassification up to 56.65 percent.

Can I just raise the detection threshold to be safe?

No. A threshold trades one error type for the other, but it cannot separate groups that genuinely overlap on the measured signal. Raising it blinds the tool; lowering it catches more honest writers. There is no value at which the tool becomes a fair exclusion rule, which is why the recommendation is to stop using it as one rather than to tune it.

Is there any defensible way to detect AI-written responses?

Not at the level of an individual response, no. What is defensible is measuring a whole batch: track lexical diversity and near-duplicate rates wave over wave and act on movement in the aggregate. That approach accuses nobody, has no disparate impact on non-native speakers, and still tells you when a question has started attracting generic answers.

Should I tell participants I am running their answers through a detector?

If you are going to act on the output in any way that affects them, yes, and that disclosure requirement is itself a good reason not to. A far better use of the same honesty is to explain up front that you want their own words and why, then ask directly and without penalty whether they used an AI tool. That gives you a usable covariate instead of a contested score.

Does this mean AI-assisted responses are not a real problem?

They are a real problem. The evidence on participant LLM use and on the resulting homogenization of open-ended text is solid. The argument here is narrower and only about the remedy: detection at the individual level is the wrong instrument for it. Fix it with question design that demands participant-specific detail, voice collection where risk is high, and corpus-level monitoring.

Related Resources

Related Articles

Attention Check Questions: How to Catch Low-Effort Survey Responses Without Annoying Real Participants

Attention check questions catch inattentive, low-effort, and fraudulent survey responses. Learn the main types, how many to use, the pitfalls, and why a conversational AI interview reduces the need for them in the first place.

Sampling Bias: Types, Examples, and How to Avoid It

Sampling bias is when some people in your population are systematically more likely to end up in your sample than others — quietly invalidating your findings. Learn the six main types, classic examples, and how to build a representative sample at scale.

Screener Accuracy: Why Most People Who Pass Your Screener Are Not Who You Wanted (2026)

A research screener is a diagnostic test. At a 5% target incidence, a screener with 90% sensitivity and 85% specificity delivers a sample that is 76% wrong. How to compute positive predictive value, measure it on your own studies, and raise it.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Fraud & Respondent Quality: How to Detect Fake and Low-Effort Responses (2026)

Between 5% and 26% of survey responses are fraudulent, and AI-generated answers now pass standard quality checks. Learn the warning signs, the detection tactics that still work, and how Koji's conversational quality gate filters bad data before it reaches your report.

Survey Response Bias: The 7 Types That Distort Your Data (and How to Reduce Them)

Response bias is the systematic distortion in how people answer research questions — from telling you what they think you want to hear, to agreeing with everything, to misremembering. This guide breaks down the seven most common response biases and how to reduce each one.

Synthetic Users in Research: Validity, Bias, and When AI Personas Are (and Aren't) Trustworthy

A research methodology guide to synthetic users — what they are, the documented bias problems (sycophancy, sign-flipping, shallow insights), the legitimate use cases, and why real AI-moderated interviews are now fast enough that the synthetic-vs-real tradeoff has fundamentally shifted.