Back to docs
Research Methods

Translating Research Questions: Why Back-Translation Fails and What to Do Instead

How to translate an interview guide or survey so every language measures the same thing: why back-translation misses problems, how the TRAPD team approach works for small teams, and how to write questions that translate well.

Short answer: do not rely on back-translation to check a translated interview guide or survey. Write the source questions so they can be translated, have the translation reviewed by people who know both the language and the research goal, test it with a few real participants in each language, and write down every decision. Survey methodologists call this team approach TRAPD: Translation, Review, Adjudication, Pretesting and Documentation. Back-translation can still be one input, but on its own it misses real problems and flags false ones.

If you run research in more than one language, translation is a measurement problem, not a copy-editing task. A question that means something slightly different in German than in English produces a German answer to a different question. The difference then shows up in your report as a "market difference" that is really a wording difference.

Why multilingual research is the normal case

Even a single-country study is often multilingual. In the United States, the Census Bureau reported that more than 1 in 5 people (22%) aged 5 and older spoke a language other than English at home in 2017 to 2021. Among Spanish speakers in the 2018 to 2022 estimates, 61.0% said they spoke English "very well", which means a large minority did not. Interviewing everyone in English quietly drops or distorts the people who would answer more fully in their first language.

Customers also respond to language directly. CSA Research surveyed 8,709 consumers in 29 countries in 2020 and found that 76% of online shoppers prefer to buy products with information in their native language, and 40% will never buy from websites in other languages. If language shapes whether people buy, it also shapes how they explain why they bought, churned or complained.

Large cross-national surveys take this seriously. The European Social Survey requires a translation for each language used as a first language by 5% or more of a country's population. Few product teams need that rigour, but the principle transfers: decide in advance which languages your audience actually needs, rather than defaulting to the one your team speaks.

Why back-translation is not enough

Back-translation means translating your questionnaire into the target language, having a second translator translate it back into the source language without seeing the original, and comparing the two source-language versions. If they match, the translation is assumed to be good.

The method is popular because it lets a team that does not speak the target language feel in control. The problem is that it checks the wrong thing. Dorothée Behr, a survey translation researcher at GESIS, studied back-translation documentation from the 2012 European Quality of Life Survey and compared it with independent reviews of the actual translations. Her conclusion, published in the International Journal of Social Research Methodology in 2017: "While back translation can uncover problems, it causes quite a number of false alarms, and even more importantly, many problems remain hidden."

There are three structural reasons:

  1. Literal translations pass easily. A word-for-word translation back-translates neatly, even when it reads awkwardly or shifts meaning for a native speaker. The translation agency cApStAn gives the example of a question about how successful the police are at catching people who commit "house burglaries". Translated literally into Slovak, the term covered houses but not flats, so the question measured something narrower than intended. The back-translation would still say "house burglaries".
  2. Register and tone disappear. Many languages distinguish a formal and informal "you". Both come back as "you" in English, so a translation that addresses participants too formally or too casually looks fine on the back-translation sheet.
  3. It tells you there is a problem, not how to fix it. Even when back-translation flags a mismatch, it gives the translator no guidance on what the right target-language wording is.

The TRAPD approach

TRAPD was developed in cross-cultural survey research, associated with Janet Harkness and adopted by the European Social Survey, which has used a team approach to translation since its first round. Here is what each step means, scaled to a product research team.

T: Translation

Two people translate independently, or each translates half the instrument. Independent versions surface different interpretations early. In a small team, one professional translator plus one bilingual colleague who knows the product is a workable minimum.

R: Review

Translators and a reviewer meet to compare versions line by line and agree on wording. The reviewer should understand what each question is trying to measure, not just the words. This is where you catch questions that are grammatically right and conceptually wrong.

A: Adjudication

Someone makes the final call where reviewers disagree. In a product team that is usually the researcher who owns the study, ideally with a native speaker present. The point is that a named person signs off, so disagreements do not stay unresolved.

P: Pretesting

Run the translated questions with a small number of real participants in each language and listen for hesitation, requests for clarification and answers that do not fit the question. A cognitive interview ("what does this question mean to you, in your own words?") is the most direct way to check that a translated question is understood as intended.

D: Documentation

Record every non-obvious decision: why a term was adapted, which examples were swapped for local ones, what the pretest changed. Six months later, when someone asks why the French numbers differ, this record tells you whether the cause could be wording.

Write questions that can be translated

Most translation problems are created in the source language. Before anything goes to a translator:

  • Remove idioms and metaphors. "Is the tool a no-brainer?" or "What moves the needle for you?" will not survive translation intact.
  • Use one term per concept. If the English version says "workspace" in one question and "account" in the next, translators will pick different words and participants will assume you mean different things.
  • Avoid double meanings. "How often do you use reports?" could mean reading or creating them. Pick one.
  • Separate examples from the question. Examples are often culture-specific ("for example, at Thanksgiving"). Mark them as replaceable so translators can choose local equivalents.
  • Keep scale labels simple. Labels such as "somewhat" and "fairly" have no exact equivalents in many languages. Fewer, clearer labels translate more reliably.
  • Add translator notes. For each question, write one line on what you are trying to learn. This lets translators prioritise meaning over form.

These are the same habits that make questions clearer for native speakers. See How to Write Unbiased Survey Questions for wording principles that apply in any language.

Adapt, do not just translate

Some questions need adaptation, not translation. Payment methods, job titles, school systems, public holidays and what counts as a "small business" vary between countries. A question can be translated perfectly and still not make sense locally. The review step should ask two questions for each item: is the translation accurate, and does the question still make sense here? If the answer to the second is no, change the question and document why. Cross-Cultural User Research covers the cultural side of this in depth, including response style differences that affect rating scales.

Analysing multilingual results

  • Compare like with like. Before you report a difference between language groups, check the translations of the questions involved. A difference on one question only is often a wording issue; a consistent difference across related questions is more likely real.
  • Code in one language, quote in the original. Agree a shared codebook, then keep original-language quotes alongside translations in the report so native speakers can check the interpretation.
  • Watch scale use. Groups can differ in how they use rating scales regardless of what they think. Open follow-up answers help you check whether a lower score reflects a real problem.

How Koji helps

Koji runs AI-moderated interviews in 28 languages, by text and, where a voice is available for the language, by voice. You choose which languages a study offers, and participants pick theirs from the list, shown in their own language. The AI interviewer then asks your questions and its follow-ups in that language.

That changes where the translation work happens, and it is worth being clear about what it does and does not replace.

  • One brief, many languages. You write the study once. The interviewer delivers it in each enabled language, so you do not maintain separate translated copies of the guide.
  • Follow-ups in the participant's language. Translated surveys only translate the fixed questions. In an AI-moderated interview, clarifications and probes also happen in the participant's language, so a participant who misunderstands a question can be guided back to it.
  • Answers that roll up together. Every structured question keeps a stable identity across languages, so answers from German, Japanese and Spanish participants to the same question land in the same chart in the report, with thematic analysis across all of them.
  • Pretest every language before launch. The Preview tab on a study's Interviews page lets you take your own interview, by text or voice, and pick the language. Previews are free and never appear in the report. This is where you do the R and P of TRAPD: ask a native-speaking colleague to run a preview in their language and flag anything that sounds wrong or means something different.

What Koji does not do is replace review and documentation. The interviewer translates your intent as you wrote it, so source questions with idioms or ambiguous terms will still cause problems. Treat the AI's rendering as the first translation in TRAPD, then review, pretest and document as you would for any translation. Large survey programmes are moving in a similar direction: the European Social Survey's draft specification for Round 13 describes a modified TRAPD that adds a machine translation step, with human review still in place.

Compared with a traditional survey tool, where each language version is a separate form you have to keep in sync, you maintain one brief and test each language directly. Compared with hiring moderators for each market, you get consistent delivery in every language without scheduling across time zones. See Multi-Language User Research for the set-up steps.

A practical checklist

  1. List the languages your audience actually needs.
  2. Clean the source questions: no idioms, one term per concept, replaceable examples, translator notes.
  3. Produce the first translation (human or AI).
  4. Review with a native speaker who understands the research goal.
  5. Have one named person adjudicate disagreements.
  6. Pretest with two or three participants per language, ideally with cognitive interview probes.
  7. Document every adaptation and why it was made.
  8. In analysis, check wording before you report a difference between language groups.

Key takeaways

  • Translation is a measurement step. A wording difference becomes a fake market difference.
  • Back-translation misses real problems and creates false alarms; use it, if at all, as one input.
  • TRAPD (Translation, Review, Adjudication, Pretesting, Documentation) is the established alternative and scales down to small teams.
  • Most translation problems start in the source questions. Write for translation from the start.
  • AI-moderated interviews remove the work of keeping separate versions in sync, but review and pretesting by native speakers still matter.

Related Resources

Sources

  • Behr, D. (2017). Assessing the use of back translation: the shortcomings of back translation as a quality testing method. International Journal of Social Research Methodology, 20(6), 573–584.
  • U.S. Census Bureau (2023). Language spoken at home, 2018–2022 American Community Survey 5-year estimates (press release, December 2023).
  • U.S. Census Bureau (2025). 2017–2021 ACS language use tables (press release, June 2025).
  • CSA Research (2020). Can't Read, Won't Buy – B2C (press release, July 2020).
  • European Social Survey. Translation methodology and Round 11 translation guidelines.
  • cApStAn. Why back translation is inadequate to assess quality in translated surveys.
  • Harkness, J. A. (2011). Translation. In Guidelines for Best Practice in Cross-Cultural Surveys. Survey Research Center, University of Michigan.

Related Articles

Cognitive Interviews: How to Test Your Survey Questions Before You Launch

A practical guide to cognitive interviewing — the pretesting technique that reveals whether your survey questions and interview guides are understood as intended. Covers think-aloud, verbal probing, sample sizing, and AI-powered approaches.

Cross-Cultural User Research: The Complete Guide for Global Product Teams

Master cross-cultural user research with frameworks for cultural adaptation, language localization, and AI-powered global insights. Avoid the bias that breaks products in new markets.

Multi-Language User Research: How to Interview Participants in Any Language

How to configure Koji to run voice and text interviews in 15+ languages — including brief localization, cross-market analysis, and synthesis best practices.

Pilot Study in User Research: How to Pre-Test Your Methodology Before Going Live (2026)

A pilot study is a small-scale rehearsal of your full research project that catches broken questions, biased prompts, and recruiting issues before they invalidate your real data. Learn when to run one, how many participants you need, what to test, and how AI-moderated platforms compress the pilot loop from weeks to hours.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

How to Write Unbiased Survey Questions: Avoiding Leading, Loaded & Double-Barreled Questions

A practical guide to question wording — the biggest hidden source of bad data. Learn to spot and fix leading, loaded, double-barreled, and assumptive questions, with real research examples and a pre-launch checklist.