Split-Ballot Experiments: How Much of Your Number Is the Question?
Write two versions of the item, randomly assign half your sample to each, and the gap is the wording effect. The technique that tells you whether your metric is a fact about customers or about your questionnaire.
A split-ballot experiment is the only way to find out how much of your result came from your customers and how much came from your question. You write two versions of the item, randomly assign half your sample to each, and the difference between them is the wording effect - measured, in your own data, with your own users. It is the single most under-used technique in product research, and it is the one that tells you whether the number you are about to put in a roadmap deck is a fact about the world or a fact about your questionnaire.
The method is old, cheap and simple. Here is Pew Research Center describing it: "We often write two versions of a question and ask half of the survey sample one version of the question and the other half the second version. Thus, we say we have two forms of the questionnaire. Respondents are assigned randomly to receive either form, so we can assume that the two groups of respondents are essentially identical. On questions where two versions are used, significant differences in the answers between the two forms tell us that the difference is a result of the way we worded the two versions."
Random assignment is what makes it decisive. The two groups are equivalent, so nothing except the wording can explain the gap.
What the gaps look like
| Experiment | Version A | Version B | Result |
|---|---|---|---|
| Military action in Iraq, January 2003 | "favor or oppose taking military action in Iraq to end Saddam Hussein's rule" | Same, plus "even if it meant that U.S. forces might suffer thousands of casualties" | 68% favor / 25% oppose becomes 43% favor / 48% oppose |
| End-of-life care, 2005 | "making it legal for doctors to give terminally ill patients the means to end their lives" | "making it legal for doctors to assist terminally ill patients in committing suicide" | 51% favor becomes 44% favor |
| Presidential priorities, January 2002 | "more important for President Bush to focus on domestic policy or foreign policy" | Same, with "foreign policy" narrowed to "the war on terrorism" | 52% domestic / 34% foreign becomes 33% domestic / 52% war on terrorism |
| Social spending | Expanding "welfare" | Expanding "assistance to the poor" | Consistently much greater support for the second across several experiments |
The Iraq item is the one to sit with. A 25-point swing, on the same day, from equivalent random halves of the same sample, caused by a subordinate clause. Both figures are real. Neither is an error.
The capstone: some numbers do not exist without their question
This is where a survey-error framework runs out of road, and it is worth being precise about why.
The total survey error budget treats every gap between your estimate and the truth as an error with a source you can name and, in principle, shrink. Paradata reads the process record to find which items are producing that error. Mode effects show that the channel is part of the measurement. All three share a premise: there is a true value out there, and better method gets you closer to it.
Split-ballot results break that premise for a specific and very common class of questions. Support for military action in Iraq in January 2003 was not "really" 68 percent, with 43 percent being the biased reading, or the reverse. Both numbers are correct answers to different questions asked of equivalent people at the same moment. There is no measurement error to reduce here, because nobody was wrong. The quantity is a joint product of the population and the instrument.
Attitudes, evaluations, priorities, stated willingness to pay and feature importance all behave this way. Behaviour mostly does not: how many times you exported a report last month has an answer that exists whether or not anyone asks. The practical rule that follows is sharp and easy to apply.
Never report an attitudinal percentage as a fact about the world without either a split-ballot estimate of how sensitive it is to wording, or the exact wording printed next to it. A number without its question is not a finding. It is half of one.
This also explains a frustration every research team eventually has: two studies about the "same" thing disagree, everyone hunts for a methodological flaw, and there is not one. Different wording, different number, both correct. The split ballot is how you convert that argument into a measurement.
Three ways to use a split ballot
1. As a diagnostic, run once. Take the one question your quarterly metric depends on, write a defensible alternative, run both for one wave, and measure the gap. If it is small, you have earned the right to stop worrying and standardise on one version forever. If it is large, you have learned that your trend line is more fragile than anyone believed - and you still standardise on one version forever, but now you know what the number is worth.
2. As permanent randomisation of an arbitrary choice. Some design decisions have no correct answer and a real effect. Scale direction is the clearest case. Pew handles it by splitting the sample: in its abortion question, "half of the sample is asked whether abortion should be legal in all cases, legal in most cases, illegal in most cases, illegal in all cases, while the other half of the sample is asked the same question with the response categories read in reverse order." The rationale is the important part: "reversing the order does not eliminate the recency effect but distributes it randomly across the population." Pew applies the same logic to the order of items in closed-ended lists, so that "no one issue appeared early or late in the list for all respondents."
That is a general and powerful move. When a design choice is arbitrary and consequential, randomising it converts a bias into variance. Bias is a permanent tilt in your estimate. Variance is noise you can quantify and average over. Trading the first for the second is almost always a good deal.
3. As a decision test. This is the version that matters most commercially. Write the two versions so that each reflects how a different stakeholder describes the problem - the version sales uses and the version engineering uses, or the optimistic framing and the sceptical one. If both versions produce the same decision, the disagreement in the room was never about the customer. If they produce different decisions, you have located the real crux, and it is a question of framing that no amount of extra sample would ever have resolved.
What to split on
| Split | Version A | Version B | What it measures |
|---|---|---|---|
| Loaded versus neutral term | "our AI assistant" | "the automated suggestions feature" | How much of your result is brand language |
| Consequence clause | "Would you use this?" | "Would you use this, if it added a step to your current workflow?" | The gap between appeal and willingness to pay a cost |
| Frame | "keeps 90% of your data clean" | "leaves 10% of your data unclean" | Gain versus loss framing sensitivity |
| Reference period | "in the last week" | "in a typical month" | How much your usage estimate depends on recall window |
| Scale direction | Best option first | Worst option first | Order and recency effects, which you then randomise permanently |
| Attribute label | "reliability" | "it works every time I need it" | Whether respondents share your vocabulary |
Three design rules keep the result interpretable.
Change exactly one thing. Two differences between forms produce a difference you cannot attribute. This is the most common way a split ballot is wasted.
Give each arm enough sample. A split ballot halves your effective n per version. As a rough guide, detecting a 10-point difference between two arms takes on the order of 200 respondents per arm at conventional confidence levels; smaller effects need substantially more. If your study has 60 respondents, a split ballot will not resolve anything statistically - but it can still be run as a qualitative probe, where the point is what people say about each version rather than the percentages. See statistical significance for the arithmetic.
Decide in advance which comparison matters, and report both arms. A split ballot analysed after the fact, with the more convenient arm reported, is worse than no split ballot, because it launders a wording choice as a finding.
Where this sits next to the tools you already have
Split-ballot testing is often confused with two adjacent practices, and the differences are practical.
Cognitive interviews test questions with a handful of people before launch and tell you why an item is misread. They are qualitative, small-n and diagnostic. Split ballots are quantitative, run at scale, and tell you how much the misreading is worth in points. Use cognitive interviews to generate the two versions; use a split ballot to price the difference between them.
Question wording rules and the framing effect tell you which constructions to avoid - leading, loaded, double-barrelled. Those rules eliminate the wording problems that are unambiguously mistakes. Split ballots handle the residual: the choices where both versions are defensible and the answer still moves. Rules cannot help you there, because neither version breaks a rule.
Question order bias covers a related family of context effects, which are also often studied with split forms but are a different design problem.
Running a split ballot in Koji
The mechanics are straightforward: create two studies that are identical except for the item under test, split your recruitment list randomly between them, and compare. Keeping everything else constant is what the random assignment is protecting, so do not also change the invitation copy, the timing or the audience.
Two things make this materially better than a classic split ballot on paper.
Structured questions keep the comparison clean. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - and a split ballot is only interpretable when both arms produce the same kind of value. Test wording on a scale or single_choice item and you get two distributions you can subtract. It is also worth noting Pew's own finding here: a 2019 study led the Center to conclude that forced-choice questions tend to yield more accurate responses than select-all-that-apply lists, especially for sensitive questions, and the Center now generally avoids select-all formats. If you are choosing between multiple_choice and a series of yes_no items for something sensitive, that is a real reason to prefer the second. See structured questions.
AI follow-ups give you the mechanism, which a classic split ballot never does. The historical limitation of split-ballot testing is that it tells you the two versions differ by nine points and is completely silent about why. In an AI-moderated interview, each arm can carry an open-ended follow-up that probes the answer the respondent just gave - "what made you say that?" - and the probes in the two arms are directly comparable. When version B respondents keep raising a cost that version A respondents never mention, you have not just measured the wording effect, you have identified the consideration the first wording suppressed. That is the difference between knowing your number is fragile and knowing what it is fragile about.
Because Koji runs the same guide in voice and text, hold the mode constant across arms as well - or at least check that the mode mix is similar - so that you are measuring a wording effect and not a wording-plus-mode effect.
Start here
Pick the single number your team argues about most. Write the version of the question a sceptic would write. Run both for one wave. Whatever the gap turns out to be, you will never again present that metric without its wording attached - and that habit alone is worth more than most methodology upgrades a research team will make this year.
Frequently asked questions
What is a split-ballot experiment?
A split-ballot experiment, also called a split-sample or split-form design, randomly assigns respondents to receive one of two versions of a question so that the difference in answers can only be caused by the wording. Because assignment is random, the two groups are equivalent, and the gap between them is a direct measurement of how much the phrasing is worth in percentage points.
How is a split ballot different from a pretest or a cognitive interview?
A cognitive interview is qualitative, uses a handful of participants, and explains why a question is misunderstood. A split ballot is quantitative, runs at full scale during the real study, and quantifies how much two defensible versions differ. The two are complements: use cognitive interviews to write the alternative version, and a split ballot to price the difference.
How large can a wording effect be?
Large enough to reverse a conclusion. In a January 2003 Pew Research Center survey, 68 percent favored and 25 percent opposed military action in Iraq; adding the clause "even if it meant that U.S. forces might suffer thousands of casualties" moved the same measurement to 43 percent in favor and 48 percent opposed. A 2005 experiment moved support for an end-of-life policy from 51 percent to 44 percent by changing "the means to end their lives" to "assist in committing suicide."
How many respondents do I need for a split ballot?
Enough in each arm to detect the difference you care about. As a rough guide, detecting a 10-point gap between two arms takes on the order of 200 respondents per arm at conventional confidence levels, and smaller effects need considerably more. Below that, run the split anyway but treat it qualitatively - what people say about each version is informative even when the percentages are not conclusive.
Should I split-test every question?
No. Splitting costs sample, and most questions do not carry a decision. Split the one or two items your decision actually rests on, and use randomisation rather than testing for arbitrary design choices such as scale direction and option order, where the goal is to spread the effect across the sample rather than to measure it.
If two wordings give different answers, which one is correct?
Usually both. For attitudes, evaluations and stated preferences, the number is a joint product of the population and the question, so there is no single true value waiting to be uncovered. The correct response is to standardise on one wording, report it alongside the number, and hold it constant over time - which is also why changing the wording of a tracked metric breaks the trend line permanently.
Related Resources
- Structured Questions Guide - the six question types, and choosing a format that makes two arms comparable
- Survey Question Wording - the wording mistakes that rules can eliminate before you get to testing
- Cognitive Interviews - the qualitative half of question testing
- The Framing Effect - why gain and loss framings of one fact produce different answers
- Total Survey Error - the error budget this technique sits inside
- Statistical Significance - working out whether the gap between two arms is real
- The AI Interviewer House Effect — the same experimental logic aimed at the asker instead of the question
Related Articles
Cognitive Interviews: How to Test Your Survey Questions Before You Launch
A practical guide to cognitive interviewing — the pretesting technique that reveals whether your survey questions and interview guides are understood as intended. Covers think-aloud, verbal probing, sample sizing, and AI-powered approaches.
The Framing Effect in Surveys and Research: How Question Wording Reverses Answers
The framing effect means the same question, worded as a gain or a loss, produces opposite answers. Learn how framing distorts surveys and interviews — and how neutral, AI-moderated question design keeps your data honest.
Mode Effects: When Letting People Choose Voice or Text Changes the Answer
Pew randomly assigned 3,003 people to phone or web and got answers that differed by up to 18 points on identical questions. Here is what that means when your respondents pick their own mode.
Paradata: What Response Time, Hesitation and Drop-Off Tell You About Your Questions
Every interview produces a record of how the answers were produced. Most teams read it to judge respondents. Read it to judge your questions instead, and you get the cheapest instrument improvement available.
Statistical Significance in Survey Research: A Plain-English Guide (2026)
A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
How to Write Unbiased Survey Questions: Avoiding Leading, Loaded & Double-Barreled Questions
A practical guide to question wording — the biggest hidden source of bad data. Learn to spot and fix leading, loaded, double-barreled, and assumptive questions, with real research examples and a pre-launch checklist.
Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)
Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.