Mode Effects: When Letting People Choose Voice or Text Changes the Answer
Pew randomly assigned 3,003 people to phone or web and got answers that differed by up to 18 points on identical questions. Here is what that means when your respondents pick their own mode.
The channel you ask through is part of the measurement, not a delivery detail. When Pew Research Center randomly assigned 3,003 people to answer the same 60 questions by phone or on the web, the two groups gave different answers on question after question - a mean difference of 5.5 percentage points, a median of five, and a maximum of 18. Nobody changed their mind. The mode changed the answer. This guide is about what that does to your research when respondents pick their own mode, which is what modern platforms - including this one - encourage them to do.
A mode effect is, in Pew's definition, "a difference in responses to a survey question attributable to the mode in which the question is administered." It is a measurement-error component, and it is one of the few that is completely invisible in your data unless you deliberately record the mode and go looking.
The evidence
Pew Research Center ran the experiment on its nationally representative American Trends Panel between July 7 and August 4, 2014, randomly assigning 3,003 respondents to a telephone survey with a live interviewer or a self-administered web questionnaire, then asking both groups the same set of 60 questions. Random assignment is what makes the study decisive: the two groups were equivalent by construction, so every difference in the answers is attributable to the channel.
| Question | Phone | Web | Difference |
|---|---|---|---|
| "Very satisfied" with family life | 62% | 44% | 18 points |
| "Very satisfied" with social life | 43% | 29% | 14 points |
| Talk to neighbours daily or a few times a week | 58% | 47% | 11 points |
| Neighbourhood is "very safe" to walk after dark | 55% | 43% | 12 points |
| Community is an "excellent" place to live | 37% | 30% | 7 points |
| "Very satisfied" with local traffic conditions | 28% | 22% | 6 points |
Twenty-one items showed a difference of at least seven points. Seven of them involved ratings of political figures, where very negative ratings were less common on the phone for all seven. Four involved intimate personal topics including life satisfaction, health and financial trouble, with more positive responses on the phone across all of them. Three concerned perceptions of discrimination.
Equally important is what did not move. Pew found no significant mode difference in how people rated their own personal happiness, or in the shares who said they had done volunteer work in the past year, called a friend or relative yesterday, or visited with family or friends yesterday. Mode effects are not a general tax on all data. They cluster on questions where an answer might reflect on you.
Three mechanisms, and only one of them is about honesty
Interviewer presence. Most of the largest differences appeared on items where social desirability could plausibly operate. Talking to a person invites a slightly better version of yourself. This is the mechanism everyone knows, and it is covered in social desirability bias and interviewer bias.
Channel and cognition. Pew is explicit that social desirability is not the whole story: "surveys require cognitive processing of words and phrases to understand a question and choose an appropriate response option," and the channel changes that processing. A complicated question with many response options is hard to hold in memory when heard aloud and easy to scan when read. Heard options also produce a recency effect, because the last option read is the easiest to remember. So a mode effect can appear on a completely neutral question purely because the question was long.
Who shows up at all. This is the mechanism product teams miss. Pew reports that even inside a carefully controlled panel experiment, "respondents who had ignored all previous survey requests were more likely to respond when they were contacted over the phone." Mode does not only change what people say; it changes which people say anything. That is why Pew closes with a sentence that belongs in every research plan: "Researchers should carefully consider the trade-offs between measurement error on the one hand and coverage and nonresponse error on the other." Mode is a line item in your total survey error budget, and it appears on both sides of the ledger with opposite signs.
The part that applies to you: self-selected mode
Standard advice for modern research platforms, including our own, is that respondents should complete a study in whichever mode suits them. That advice is right about participation - see voice vs text for the trade-offs of each channel - and it has a consequence that the advice usually leaves out.
If respondents choose their mode, mode is no longer a design variable. It is a respondent characteristic, and it is correlated with everything that makes someone prefer talking to typing: seniority, whether they are at a desk, whether they are commuting, how comfortable they are in the interview language, whether they are in an open-plan office, how much time they have. Which means:
Any comparison between two groups with different mode mixes is partly a comparison of modes.
That is the whole risk, and it shows up in three concrete places.
| Situation | How it bites | Severity |
|---|---|---|
| Tracking a metric across waves | Your voice share drifts from 30% to 55% between quarters because of a UX change or a different audience; the metric moves and nothing in the world changed | High - trend data is the most mode-sensitive artefact there is |
| Comparing segments | Field staff answer by voice, head-office staff by text; you attribute the gap to role | High, and usually invisible |
| Comparing against an external benchmark | Your conversational voice study reports higher satisfaction than an industry web benchmark | High - almost always partly mode |
| One-off discovery research | You are looking for themes and mechanisms, not point estimates | Low - mode differences barely matter here |
| A/B comparison inside one study | Both arms have the same mode mix, so mode cancels | Low, if the mix really is the same |
The tracking case is the one that hurts most, because a mode-driven shift is indistinguishable from a real shift in the data itself. Pew, which has more resources for this problem than any product team, describes the challenge plainly as "how to track trends in responses over time when the mode of interview has changed," and notes that researchers are still "developing methods for combining data collected from different modes so that disruption to long-standing trend data is minimized."
What to do about it
You do not need to abandon multi-mode research. You need five habits.
1. Record mode on every response, always. This is the entire foundation. A mode column costs nothing and makes every check below possible. Without it, the confound exists and is permanently unmeasurable.
2. Report the mode mix next to the number. "NPS 41, 62 percent text" is a defensible statistic. "NPS 41" is a number whose meaning depends on a variable you did not disclose.
3. Hold the mix roughly constant on anything you track. For a brand tracking study or a quarterly satisfaction metric, treat a large shift in mode mix the same way you would treat a change in the sampling frame: as a break in the series that must be flagged. If the mix moves more than about 10 points between waves, say so in the report.
4. Apply unified mode design. Write the instrument so the stimulus is as close to identical as possible in both channels. In practice: keep response options to four or five so they survive being heard aloud; avoid long stems; avoid asking respondents to hold a list in memory; use the same wording in both modes rather than a "spoken version" and a "written version." A question that has to be rewritten for voice is a question that will produce a mode effect.
5. Force one mode when the number is the deliverable. If the output is a point estimate that will be compared to something - a benchmark, a target, last quarter - run a single mode and say which. Save mixed mode for studies where the output is understanding.
A useful sanity check when you already have mixed data: split your headline metric by mode. If the gap is small relative to your decision threshold, proceed and note it. If the gap is large, you have learned something important before you shipped a wrong conclusion, and the correct next step is to compare within mode rather than to average across them.
How Koji helps, and where it does not
Most research stacks make mode effects worse by accident: the voice study happens in a video call with a moderator and a discussion guide, the text study happens in a form builder with different wording, and the two datasets get merged in a slide. That is two different instruments as well as two different modes, and the two are then permanently inseparable.
Koji is unusually well positioned on this specific problem for a structural reason: the same study, with the same interview guide and the same questions, runs in both voice and text.
- The instrument is genuinely unified. Unified mode design is normally an aspiration that survives until someone rewrites a question for the phone script. When both modes are driven by one guide, the stimulus really is the same, which removes the largest and most avoidable source of mode difference.
- Mode is recorded on every interview, so the checks above are a filter rather than a research project.
- Structured questions travel across modes. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and using structured types for anything you intend to trend gives you a comparable value in both channels rather than a paraphrase. Long option lists remain the enemy in voice, which is why keeping choice sets short is a design rule and not a preference. See structured questions.
- AI follow-ups reduce the cognitive-processing mechanism. A respondent who did not understand a heard question gets a clarification in the moment, in both modes, instead of guessing. That removes part of the comprehension gap that drives non-sensitive mode effects.
- Mode has a cost dimension too. In Koji, a text conversation consumes 1 credit and a voice conversation consumes 3, so forcing a single mode for a tracking study is also a budget decision you can make deliberately rather than by accident.
What Koji does not do is make mode effects disappear. Removing the human moderator removes interviewer-driven social desirability, which is a real and well-documented reduction, but voice and text are still different channels reaching different people in different moments. The honest position is that Koji makes mode a variable you can see and control, rather than one that quietly rides along inside your trend line.
An honest limitation
Pew's experiment tested public-opinion questions - satisfaction, discrimination, ratings of political figures. Product research asks a lot of behavioural questions ("how many times did you export a report last month") where mode effects are smaller, exactly as Pew found for reports of yesterday activities and volunteering. Treat the 18-point family-life gap as an upper bound for an attitudinal, self-reflective item, not as a forecast for your usage questions. The mechanism transfers; the magnitude has to be checked in your own data, which is precisely what recording the mode column lets you do.
Frequently asked questions
What is a mode effect in research?
A mode effect is a difference in responses to the same question that is caused by the channel used to ask it - voice versus text, interviewer-administered versus self-administered, phone versus web. It is a measurement-error component, and because the question wording is identical, it is invisible in your dataset unless the mode of each response is recorded.
How big are mode effects in practice?
In Pew Research Center's randomised experiment with 3,003 respondents across 60 questions, the mean difference between telephone and web was 5.5 percentage points and the median was five, with a range from zero to 18 points. The largest gaps appeared on self-reflective items such as satisfaction with family life, where 62 percent said "very satisfied" by phone against 44 percent on the web. Purely behavioural questions often showed no difference at all.
Is it a problem if respondents choose their own mode?
For discovery research, rarely. For any number you intend to compare - across waves, across segments, or against a benchmark - yes, because self-selected mode is correlated with respondent characteristics, so a mode difference and a segment difference become impossible to separate. The fix is not to stop offering choice; it is to record mode, report the mix, and check whether your headline metric differs by mode.
How do I stop mode effects from breaking a tracking study?
Hold the mode mix approximately constant across waves and treat a large shift the way you would treat a change of sampling frame - as a documented break in the series. If the study exists to produce one comparable number, run it in a single mode. And apply unified mode design so the instrument itself is identical in both channels rather than rewritten for each.
Does removing the human interviewer eliminate mode effects?
It removes one important mechanism - the social desirability pressure created by another person listening - which is why self-administered modes consistently produce more candid answers on sensitive topics. It does not remove the others. Aural and visual presentation still impose different cognitive loads, and different people still choose different channels, so mode remains a variable worth recording.
Which questions are most vulnerable to mode effects?
Questions where the answer reflects on the respondent, and questions that are hard to process by ear. Satisfaction with your own life, ratings of people, sensitive behaviours and perceptions of social problems moved most in Pew's data. Long question stems and lists of more than four or five response options are the second family, because they are easy to scan and hard to hold in memory.
Related Resources
- Structured Questions Guide - the six question types and how to keep them comparable across channels
- Voice vs Text Interviews - when each mode wins, and how to decide
- Total Survey Error - where mode sits in the full error budget
- Social Desirability Bias - the mechanism behind the largest mode gaps
- Interviewer Bias - how moderator presence distorts answers
- Brand Tracking Study Guide - running a wave-over-wave metric without breaking the series
- The AI Interviewer House Effect — the other confound you cannot let participants self-select into
Related Articles
Brand Tracking Studies: How to Measure Brand Health Over Time (2026)
A complete guide to brand tracking studies — what to measure, how often to run them, sample size, and how AI-native platforms make continuous brand tracking affordable for the first time.
Interviewer Bias: How Moderators Distort Research (and How AI Removes the Variance)
Interviewer bias is the distortion caused by a moderator's wording, reactions, expectations, and characteristics. Learn the types, the evidence, mitigation techniques, and why an AI interviewer eliminates interviewer variance.
Paradata: What Response Time, Hesitation and Drop-Off Tell You About Your Questions
Every interview produces a record of how the answers were produced. Most teams read it to judge respondents. Read it to judge your questions instead, and you get the cheapest instrument improvement available.
Social Desirability Bias: What It Is and How to Eliminate It in Research
Social desirability bias makes people tell you what sounds good instead of what is true. Learn what causes it, why it quietly wrecks product decisions, and the seven evidence-based ways to reduce it — including why AI-moderated interviews get more honest answers.
Split-Ballot Experiments: How Much of Your Number Is the Question?
Write two versions of the item, randomly assign half your sample to each, and the gap is the wording effect. The technique that tells you whether your metric is a fact about customers or about your questionnaire.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)
Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.
Voice vs Text Interview: When to Use Each Mode
Choosing between voice and text mode for your AI interview? This guide breaks down response depth, completion rate, audience fit, and cost — plus a decision matrix that tells you which mode wins for each research scenario.