{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-10-01T10:10:40.893Z"},"content":[{"type":"documentation","id":"1f42f834-a01c-422c-882a-d019a2f3ec71","slug":"interpreter-mediated-interviews-error","title":"What a Human Interpreter Changes in a Research Interview (2026)","url":"https://www.koji.so/docs/interpreter-mediated-interviews-error","summary":"Human interpreters introduce a measurable error rate in research interviews. Flores et al. (Pediatrics, 2003) found 396 errors across 13 encounters, a mean of 31 each, with 63 percent carrying potential clinical consequence, distributed as omission 52 percent, false fluency 16 percent, substitution 13 percent, editorialization 10 percent, addition 8 percent. Flores et al. (Annals of Emergency Medicine, 2012) compared arrangements across 57 encounters and 1,884 errors: 12 percent of errors carried potential consequence with professional interpreters, 22 percent with ad hoc interpreters and 20 percent with no interpreter, so an untrained bilingual helper scored no better than none. Training hours predicted accuracy while years of experience did not. Omission is the most dangerous class for research because it leaves no artifact, and research has no downstream outcome to catch the error. The structural fix is interviewing in the participant's own language with the original-language quote retained alongside an English code label.","content":"When you cannot speak a participant's language, the obvious fix is to put a bilingual person in the middle. Medicine has studied what that actually does to the record, in transcript-level detail, and the findings should change how product teams run research in non-native markets. The two most important ones are counterintuitive: an untrained bilingual helper can produce a worse record than having no interpreter at all, and the most frequent error is the one that leaves no trace.\n\n## The short answer\n\nA human relay is not a neutral pipe. It introduces a measurable error rate, dominated by omission, and the errors that matter cluster in untrained interpreters - which is exactly what a helpful bilingual colleague is. If you must use one, use a trained interpreter, brief them that fidelity beats fluency, and keep the original-language audio. If you can interview in the participant's own language without a relay, do that instead, because it removes the error source rather than managing it.\n\n## The measured error rate\n\nFlores and colleagues (Pediatrics, 2003; volume 111, pages 6-14) audiotaped and transcribed pediatric encounters in which a Spanish interpreter was used. Thirteen encounters produced 474 pages of transcripts, and they catalogued 396 interpreter errors - a mean of 31 per encounter. Sixty-three percent of those errors had potential clinical consequences, a mean of 19 per encounter.\n\nTwo details from that study are worth sitting with. First, the people doing the interpreting: professional hospital interpreters covered 6 of the encounters, and the ad hoc interpreters included nurses, social workers, and an 11-year-old sibling. Second, errors committed by the ad hoc interpreters were significantly more likely to carry potential clinical consequence than those from hospital interpreters, 77 percent versus 53 percent.\n\nThe documented examples make the abstract category concrete. They include omitting questions about drug allergies, omitting dosing instructions, and - the one that should alarm any researcher - instructing a mother not to answer personal questions. That last one is an interpreter overriding the protocol rather than conveying it.\n\n## The finding that should change your default\n\nA larger follow-up compared the three arrangements directly. Flores, Abreu, Barone, Bachur and Lin (Annals of Emergency Medicine, 2012; volume 60, pages 545-53) analysed audiotaped emergency department visits over 30 months in the two largest pediatric emergency departments in Massachusetts. They report: \"The 57 encounters included 20 with professional interpreters, 27 with ad hoc interpreters, and 10 with no interpreters; 1,884 interpreter errors were noted, and 18% had potential clinical consequences.\"\n\nThe proportion of errors carrying potential consequence was 12 percent with professional interpreters, 22 percent with ad hoc interpreters, and 20 percent with no interpreter at all. Read the middle and last figures together, because that is the uncomfortable result: **the ad hoc arrangement scored no better than having no interpreter.** An untrained bilingual helper is not a partial solution on the way to a good one. On this measure it was the worst of the three options, presumably because it produces a fluent, confident record that nobody thinks to doubt, whereas an encounter with no interpreter is visibly compromised and gets treated with appropriate suspicion.\n\nThe authors conclude: \"Professional interpreters result in a significantly lower likelihood of errors of potential consequence than ad hoc and no interpreters.\"\n\nThere is a second finding with an immediate practical edge. Among the professional interpreters, previous hours of interpreter training were significantly associated with error numbers and consequences, but years of experience were not. Interpreters with at least 100 hours of training had a median of 12 errors versus 33 for those with less training, and committed 2 percent versus 12 percent errors of potential consequence. **Training predicted accuracy; experience did not.** So \"she has been here ten years and she speaks Portuguese\" is precisely the wrong credential to select on, and it is the one most teams use.\n\n## The five error types, translated\n\nThe 2003 study classified errors into five types. The distribution was omission 52 percent, false fluency 16 percent, substitution 13 percent, editorialization 10 percent, and addition 8 percent. Each has a specific research consequence.\n\n| Error type | What it is | What it costs your research |\n| --- | --- | --- |\n| Omission | Content simply not conveyed | The dominant type, and invisible. A dropped qualifier or an unrelayed follow-up leaves no artifact to audit |\n| False fluency | A word or phrase that does not exist or is wrong | Creates a confident-sounding finding with nothing behind it |\n| Substitution | Replacing a term with a different one | Quietly changes the category a response gets coded into |\n| Editorialization | The interpreter inserting their own view | The interpreter becomes a second respondent whose answers you cannot separate out |\n| Addition | Content introduced that nobody said | Manufactures evidence, and it will survive every quote check because the transcript contains it |\n\nOmission at 52 percent deserves the most attention precisely because it is the hardest to catch. A substitution or an addition puts something wrong into the record, where a bilingual reviewer can find it. An omission leaves the record shorter and perfectly coherent. Nothing in the English transcript indicates that a sentence was ever said.\n\n## Why this is worse in research than in medicine\n\nA clinical encounter has a downstream check. The patient either improves or does not, and a serious interpretation error eventually surfaces as a clinical event. That feedback loop is slow and expensive, but it exists.\n\nResearch has no such loop. The transcript is not an input to an outcome you later observe; the transcript *is* the output. If a participant's actual motivation was omitted and a plausible nearby motivation was conveyed instead, that becomes a theme, then a roadmap item, and nothing downstream ever contradicts it. Every quality check you run will pass, because every check operates on the English record, and the English record is internally consistent. This is the same structural problem as any summary that you cannot re-derive from source: the error class your checks are built to catch is not the error class you have.\n\n## Your ad hoc interpreters, named\n\nTeams rarely think they are using an ad hoc interpreter, because the category sounds like an emergency improvisation. In practice it covers almost every arrangement product teams actually use:\n\n- The bilingual account manager who joined the call to help.\n- A sales engineer in the region, who also has a commercial relationship with the participant.\n- A colleague the participant brought along.\n- A family member.\n- A generalist from a vendor who was not briefed on research interpreting.\n\nThe account manager case compounds two problems at once. They are untrained as an interpreter, and they have an interest in how the account is characterised. The editorialization rate is not going to be 10 percent.\n\n## If you must use a human interpreter\n\n- **Select on training, not on years or fluency.** The evidence points at hours of interpreter training specifically.\n- **Brief them that fidelity beats smoothness.** Business and clinical interpreting optimise for a successful interaction. Research interpreting optimises for an accurate record, including the hesitations, the self-corrections, and the parts that do not make sense. Tell them explicitly that tidying is a defect.\n- **Use consecutive, first-person interpreting**, and have them flag rather than resolve ambiguity.\n- **Keep the original-language audio and transcript**, not just the English. If you discard the original you have destroyed the only thing that could ever detect an omission.\n- **Spot-check a sample.** Have a second bilingual reviewer compare a handful of segments against the original audio. You are not auditing the interpreter's character; you are measuring your own error rate.\n- **Never let a commercially interested party interpret.** This is not a competence question.\n\n## How Koji handles this\n\nThe cleanest answer to a relay error is not to manage the relay but to remove it. Koji's AI interviewer converses directly with participants in their own language, in voice or text, so there is no third party between the participant and the record. The error categories above describe what happens in a human hand-off; with no hand-off, they have nothing to act on.\n\nThe design detail that matters most for analysis is what Koji keeps. When Koji codes an open-ended answer, the theme label is produced in English so that findings stay comparable across markets, while the supporting quote is retained in the participant's original language, linked to the exact transcript messages it came from. You get a comparable code and the untranslated evidence behind it, in the same record. That combination is what makes omission detectable at all: a reviewer can go to the original words rather than inspecting an English summary for a gap that, by definition, is not there.\n\nKoji's structured questions remove the relay from the quantitative side entirely. A `scale` rating, a `single_choice` selection, a `ranking` order and a `yes_no` answer are captured as values rather than as prose, so a cross-market comparison on those items never passes through anybody's translation. Those four plus open_ended and multiple_choice are the six question types available, and in multilingual work the structured ones are disproportionately valuable.\n\nThere is also a quieter benefit in probing depth. With a human interpreter, every follow-up costs two extra conversational turns, so follow-ups get rationed and non-English interviews end up systematically shallower than English ones - a bias that never appears in any report. Koji's AI interviewer generates follow-ups in the participant's own language at a configurable depth applied to every interview, so probing does not silently degrade by market. For setup, including the current language list and brief localisation, see the multilingual research guide below.\n\n## Common mistakes\n\n- **Treating a bilingual colleague as equivalent to a trained interpreter.** The measured gap between ad hoc and professional is large, and ad hoc did not beat having no interpreter at all.\n- **Selecting on years of experience.** Training hours predicted accuracy in the data; experience did not.\n- **Discarding the original-language recording.** This permanently removes your ability to detect the most common error type.\n- **Letting an account owner interpret.** Untrained plus commercially interested is the worst available combination.\n- **Assuming a clean English transcript means a clean interview.** Omissions produce transcripts that are shorter and completely coherent.\n- **Comparing scores across languages without checking the instrument.** Interpretation fidelity is a separate problem from whether a scale means the same thing in two languages.\n\n## Frequently asked questions\n\n### How many errors does a human interpreter actually introduce?\n\nIn the Flores 2003 pediatric study, 13 encounters yielded 396 interpreter errors, a mean of 31 per encounter, and 63 percent of those errors had potential clinical consequences. The larger 2012 study across 57 encounters catalogued 1,884 errors, of which 18 percent had potential clinical consequences.\n\n### Is a bilingual colleague better than no interpreter at all?\n\nNot according to the measured data. In the 2012 study the proportion of errors with potential consequence was 22 percent for ad hoc interpreters and 20 percent with no interpreter, compared with 12 percent for professionals. An untrained helper produces a fluent record that is harder to be appropriately sceptical about.\n\n### What should I look for when hiring a research interpreter?\n\nHours of formal interpreter training, not years of experience or apparent fluency. In the 2012 data, interpreters with at least 100 hours of training had a median of 12 errors versus 33, and 2 percent versus 12 percent errors of potential consequence, while years of experience showed no significant association.\n\n### Which interpretation error is most dangerous for research?\n\nOmission, for two reasons: it was the most common type at 52 percent of errors, and it leaves no trace. A substitution or addition puts something checkable into the transcript, whereas an omission produces a shorter record that reads as complete.\n\n### Why is interpreter error a bigger problem in research than in clinical care?\n\nBecause research has no downstream outcome to catch it. In medicine a serious error eventually surfaces in the patient's course. In research the transcript is the deliverable, so an omitted motivation becomes a theme and then a roadmap item, and every quality check still passes because all of them read the English record.\n\n### Can I avoid the problem by interviewing in the participant's own language?\n\nYes, and that is the structural fix rather than a mitigation. An AI interviewer that converses directly in the participant's language removes the hand-off where these errors occur, and keeping the original-language quote alongside an English code label preserves your ability to verify any finding against what was actually said.\n\n## Related Resources\n\n- [Multi-Language User Research](/docs/multilingual-research-guide) - how to configure and run studies in other languages, including brief localisation\n- [Cross-Cultural User Research](/docs/cross-cultural-user-research) - the cultural layer that sits on top of the language layer\n- [Measurement Invariance](/docs/measurement-invariance-comparing-groups) - why a translated scale may still not be comparable across markets\n- [Hearsay in Product Research](/docs/hearsay-secondhand-customer-evidence) - the related problem of a claim that reached you through someone else\n- [Um, Uh, and the False Start](/docs/speech-disfluency-interview-transcripts) - why tidying a transcript deletes data, which is what interpreters are trained to do\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types, and why structured items need no translation\n","category":"Interview Techniques","lastModified":"2026-10-01T03:42:44.406513+00:00","metaTitle":"Interpreter Error in Research Interviews","metaDescription":"Measured interpreter error: omission dominates, and an untrained bilingual helper scored no better than using no interpreter at all.","keywords":["interpreter mediated interviews","ad hoc interpreter error","multilingual interview accuracy","translated interview research","interpreter error taxonomy","research interpreting fidelity"],"aiSummary":"Human interpreters introduce a measurable error rate in research interviews. Flores et al. (Pediatrics, 2003) found 396 errors across 13 encounters, a mean of 31 each, with 63 percent carrying potential clinical consequence, distributed as omission 52 percent, false fluency 16 percent, substitution 13 percent, editorialization 10 percent, addition 8 percent. Flores et al. (Annals of Emergency Medicine, 2012) compared arrangements across 57 encounters and 1,884 errors: 12 percent of errors carried potential consequence with professional interpreters, 22 percent with ad hoc interpreters and 20 percent with no interpreter, so an untrained bilingual helper scored no better than none. Training hours predicted accuracy while years of experience did not. Omission is the most dangerous class for research because it leaves no artifact, and research has no downstream outcome to catch the error. The structural fix is interviewing in the participant's own language with the original-language quote retained alongside an English code label.","aiDifficulty":"intermediate","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}