Back to docs
Interview Techniques

What a Human Interpreter Changes in a Research Interview (2026)

Medical research has measured interpreter error precisely. The headline: an untrained bilingual helper can be worse than no interpreter at all, and omission is the most common error.

When you cannot speak a participant's language, the obvious fix is to put a bilingual person in the middle. Medicine has studied what that actually does to the record, in transcript-level detail, and the findings should change how product teams run research in non-native markets. The two most important ones are counterintuitive: an untrained bilingual helper can produce a worse record than having no interpreter at all, and the most frequent error is the one that leaves no trace.

The short answer

A human relay is not a neutral pipe. It introduces a measurable error rate, dominated by omission, and the errors that matter cluster in untrained interpreters - which is exactly what a helpful bilingual colleague is. If you must use one, use a trained interpreter, brief them that fidelity beats fluency, and keep the original-language audio. If you can interview in the participant's own language without a relay, do that instead, because it removes the error source rather than managing it.

The measured error rate

Flores and colleagues (Pediatrics, 2003; volume 111, pages 6-14) audiotaped and transcribed pediatric encounters in which a Spanish interpreter was used. Thirteen encounters produced 474 pages of transcripts, and they catalogued 396 interpreter errors - a mean of 31 per encounter. Sixty-three percent of those errors had potential clinical consequences, a mean of 19 per encounter.

Two details from that study are worth sitting with. First, the people doing the interpreting: professional hospital interpreters covered 6 of the encounters, and the ad hoc interpreters included nurses, social workers, and an 11-year-old sibling. Second, errors committed by the ad hoc interpreters were significantly more likely to carry potential clinical consequence than those from hospital interpreters, 77 percent versus 53 percent.

The documented examples make the abstract category concrete. They include omitting questions about drug allergies, omitting dosing instructions, and - the one that should alarm any researcher - instructing a mother not to answer personal questions. That last one is an interpreter overriding the protocol rather than conveying it.

The finding that should change your default

A larger follow-up compared the three arrangements directly. Flores, Abreu, Barone, Bachur and Lin (Annals of Emergency Medicine, 2012; volume 60, pages 545-53) analysed audiotaped emergency department visits over 30 months in the two largest pediatric emergency departments in Massachusetts. They report: "The 57 encounters included 20 with professional interpreters, 27 with ad hoc interpreters, and 10 with no interpreters; 1,884 interpreter errors were noted, and 18% had potential clinical consequences."

The proportion of errors carrying potential consequence was 12 percent with professional interpreters, 22 percent with ad hoc interpreters, and 20 percent with no interpreter at all. Read the middle and last figures together, because that is the uncomfortable result: the ad hoc arrangement scored no better than having no interpreter. An untrained bilingual helper is not a partial solution on the way to a good one. On this measure it was the worst of the three options, presumably because it produces a fluent, confident record that nobody thinks to doubt, whereas an encounter with no interpreter is visibly compromised and gets treated with appropriate suspicion.

The authors conclude: "Professional interpreters result in a significantly lower likelihood of errors of potential consequence than ad hoc and no interpreters."

There is a second finding with an immediate practical edge. Among the professional interpreters, previous hours of interpreter training were significantly associated with error numbers and consequences, but years of experience were not. Interpreters with at least 100 hours of training had a median of 12 errors versus 33 for those with less training, and committed 2 percent versus 12 percent errors of potential consequence. Training predicted accuracy; experience did not. So "she has been here ten years and she speaks Portuguese" is precisely the wrong credential to select on, and it is the one most teams use.

The five error types, translated

The 2003 study classified errors into five types. The distribution was omission 52 percent, false fluency 16 percent, substitution 13 percent, editorialization 10 percent, and addition 8 percent. Each has a specific research consequence.

Error typeWhat it isWhat it costs your research
OmissionContent simply not conveyedThe dominant type, and invisible. A dropped qualifier or an unrelayed follow-up leaves no artifact to audit
False fluencyA word or phrase that does not exist or is wrongCreates a confident-sounding finding with nothing behind it
SubstitutionReplacing a term with a different oneQuietly changes the category a response gets coded into
EditorializationThe interpreter inserting their own viewThe interpreter becomes a second respondent whose answers you cannot separate out
AdditionContent introduced that nobody saidManufactures evidence, and it will survive every quote check because the transcript contains it

Omission at 52 percent deserves the most attention precisely because it is the hardest to catch. A substitution or an addition puts something wrong into the record, where a bilingual reviewer can find it. An omission leaves the record shorter and perfectly coherent. Nothing in the English transcript indicates that a sentence was ever said.

Why this is worse in research than in medicine

A clinical encounter has a downstream check. The patient either improves or does not, and a serious interpretation error eventually surfaces as a clinical event. That feedback loop is slow and expensive, but it exists.

Research has no such loop. The transcript is not an input to an outcome you later observe; the transcript is the output. If a participant's actual motivation was omitted and a plausible nearby motivation was conveyed instead, that becomes a theme, then a roadmap item, and nothing downstream ever contradicts it. Every quality check you run will pass, because every check operates on the English record, and the English record is internally consistent. This is the same structural problem as any summary that you cannot re-derive from source: the error class your checks are built to catch is not the error class you have.

Your ad hoc interpreters, named

Teams rarely think they are using an ad hoc interpreter, because the category sounds like an emergency improvisation. In practice it covers almost every arrangement product teams actually use:

  • The bilingual account manager who joined the call to help.
  • A sales engineer in the region, who also has a commercial relationship with the participant.
  • A colleague the participant brought along.
  • A family member.
  • A generalist from a vendor who was not briefed on research interpreting.

The account manager case compounds two problems at once. They are untrained as an interpreter, and they have an interest in how the account is characterised. The editorialization rate is not going to be 10 percent.

If you must use a human interpreter

  • Select on training, not on years or fluency. The evidence points at hours of interpreter training specifically.
  • Brief them that fidelity beats smoothness. Business and clinical interpreting optimise for a successful interaction. Research interpreting optimises for an accurate record, including the hesitations, the self-corrections, and the parts that do not make sense. Tell them explicitly that tidying is a defect.
  • Use consecutive, first-person interpreting, and have them flag rather than resolve ambiguity.
  • Keep the original-language audio and transcript, not just the English. If you discard the original you have destroyed the only thing that could ever detect an omission.
  • Spot-check a sample. Have a second bilingual reviewer compare a handful of segments against the original audio. You are not auditing the interpreter's character; you are measuring your own error rate.
  • Never let a commercially interested party interpret. This is not a competence question.

How Koji handles this

The cleanest answer to a relay error is not to manage the relay but to remove it. Koji's AI interviewer converses directly with participants in their own language, in voice or text, so there is no third party between the participant and the record. The error categories above describe what happens in a human hand-off; with no hand-off, they have nothing to act on.

The design detail that matters most for analysis is what Koji keeps. When Koji codes an open-ended answer, the theme label is produced in English so that findings stay comparable across markets, while the supporting quote is retained in the participant's original language, linked to the exact transcript messages it came from. You get a comparable code and the untranslated evidence behind it, in the same record. That combination is what makes omission detectable at all: a reviewer can go to the original words rather than inspecting an English summary for a gap that, by definition, is not there.

Koji's structured questions remove the relay from the quantitative side entirely. A scale rating, a single_choice selection, a ranking order and a yes_no answer are captured as values rather than as prose, so a cross-market comparison on those items never passes through anybody's translation. Those four plus open_ended and multiple_choice are the six question types available, and in multilingual work the structured ones are disproportionately valuable.

There is also a quieter benefit in probing depth. With a human interpreter, every follow-up costs two extra conversational turns, so follow-ups get rationed and non-English interviews end up systematically shallower than English ones - a bias that never appears in any report. Koji's AI interviewer generates follow-ups in the participant's own language at a configurable depth applied to every interview, so probing does not silently degrade by market. For setup, including the current language list and brief localisation, see the multilingual research guide below.

Common mistakes

  • Treating a bilingual colleague as equivalent to a trained interpreter. The measured gap between ad hoc and professional is large, and ad hoc did not beat having no interpreter at all.
  • Selecting on years of experience. Training hours predicted accuracy in the data; experience did not.
  • Discarding the original-language recording. This permanently removes your ability to detect the most common error type.
  • Letting an account owner interpret. Untrained plus commercially interested is the worst available combination.
  • Assuming a clean English transcript means a clean interview. Omissions produce transcripts that are shorter and completely coherent.
  • Comparing scores across languages without checking the instrument. Interpretation fidelity is a separate problem from whether a scale means the same thing in two languages.

Frequently asked questions

How many errors does a human interpreter actually introduce?

In the Flores 2003 pediatric study, 13 encounters yielded 396 interpreter errors, a mean of 31 per encounter, and 63 percent of those errors had potential clinical consequences. The larger 2012 study across 57 encounters catalogued 1,884 errors, of which 18 percent had potential clinical consequences.

Is a bilingual colleague better than no interpreter at all?

Not according to the measured data. In the 2012 study the proportion of errors with potential consequence was 22 percent for ad hoc interpreters and 20 percent with no interpreter, compared with 12 percent for professionals. An untrained helper produces a fluent record that is harder to be appropriately sceptical about.

What should I look for when hiring a research interpreter?

Hours of formal interpreter training, not years of experience or apparent fluency. In the 2012 data, interpreters with at least 100 hours of training had a median of 12 errors versus 33, and 2 percent versus 12 percent errors of potential consequence, while years of experience showed no significant association.

Which interpretation error is most dangerous for research?

Omission, for two reasons: it was the most common type at 52 percent of errors, and it leaves no trace. A substitution or addition puts something checkable into the transcript, whereas an omission produces a shorter record that reads as complete.

Why is interpreter error a bigger problem in research than in clinical care?

Because research has no downstream outcome to catch it. In medicine a serious error eventually surfaces in the patient's course. In research the transcript is the deliverable, so an omitted motivation becomes a theme and then a roadmap item, and every quality check still passes because all of them read the English record.

Can I avoid the problem by interviewing in the participant's own language?

Yes, and that is the structural fix rather than a mitigation. An AI interviewer that converses directly in the participant's language removes the hand-off where these errors occur, and keeping the original-language quote alongside an English code label preserves your ability to verify any finding against what was actually said.

Related Resources

Related Articles

Cross-Cultural User Research: The Complete Guide for Global Product Teams

Master cross-cultural user research with frameworks for cultural adaptation, language localization, and AI-powered global insights. Avoid the bias that breaks products in new markets.

Hearsay in Product Research: Why a Relayed Customer Claim Is Not Customer Evidence

Most product decisions rest on relayed claims about what customers want. Borrow the law of evidence's hearsay rule to grade every claim, and promote the ones that matter to first-hand evidence in 48 hours.

Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)

Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.

Multi-Language User Research: How to Interview Participants in Any Language

How to configure Koji to run voice and text interviews in 15+ languages — including brief localization, cross-market analysis, and synthesis best practices.

Um, Uh, and the False Start: What Transcript Cleanup Deletes (2026)

Filled pauses are not noise. See what um and uh signal, why listeners benefit from them, and why your transcription tool removes them by default.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.