Back to docs
Research Methods

You Changed the Question: How to Bridge a Trend Across an Instrument Change (2026)

Rewriting a tracking question breaks your trend. Run both versions in one overlap wave, then bridge the series with the measured gap or break it honestly.

If you change a tracking question and keep plotting one line, you have published a change in your instrument as a change in your customers. The fix is not to argue about which wording is better. It is to run both versions in the same wave, measure the gap between them, and then either bridge the series with that gap or break the series honestly and say so.

Every long-running research programme eventually faces this. The satisfaction item is badly worded, the scale labels are inconsistent, the question asks about a feature that no longer exists. Rewriting it is correct. The mistake comes next, when the new number is appended to the old chart and the step change is read as customer behaviour.

The failure, stated plainly

A quarterly tracker asks the same satisfaction item for two years. In Q1 you rewrite it. The number jumps from 62 to 70 and the deck reports an eight-point improvement.

But you changed two things at once: the quarter and the instrument. The eight points are some unknown mixture of real movement and wording effect, and nothing in the data as collected can separate them. There is no statistical repair after the fact. The information needed to split those eight points was available only during the wave in which both versions could have been fielded, and you did not field both.

This is a different failure from the two it is usually confused with, and the distinction decides what you do:

  • A split-ballot experiment measures how much of a number comes from its wording. That is the measurement tool. This article is about what to do with your historical trend once you know the gap, which a split-ballot by itself does not tell you.
  • A false trend from the measurement interval is aliasing: the instrument never changed, but the sampling rhythm manufactured a pattern. Here the rhythm is fine and the instrument moved.

How big is the gap you are ignoring?

Large enough to swamp the effects most trackers are built to detect. Pew Research Center states the general principle directly: "Even small wording differences can substantially affect the answers people provide."

Their published examples show the size. In a 2005 Pew Research Center survey, 51 percent of respondents said they favoured making it legal for doctors to give terminally ill patients the means to end their lives, while only 44 percent favoured making it legal for doctors to assist terminally ill patients in committing suicide. Those two phrasings describe the same policy, and they differ by seven points.

Question order does the same work. Asked whether they were satisfied with the way things were going in the country immediately after a presidential approval question, 88 percent said they were dissatisfied, compared with 78 percent without that prior context: a ten-point swing from context alone, with the item itself untouched.

Response options matter just as much. When the economy was explicitly offered as an answer, 58 percent chose it, against only 35 percent who volunteered it in the open-ended version. If you convert an open_ended item to a single_choice item, you should expect a shift of that order, and you should not call it a change in priorities.

Seven to twenty-three points. Most product trackers are looking for movements of two or three.

What the statistical agencies actually do

National statistics offices face this problem with consequences, and their solution is worth copying because it is unglamorous and it works: field both instruments at once, in the same period, on randomly assigned subsamples.

When the US Census Bureau redesigned the income questions in the Current Population Survey Annual Social and Economic Supplement for 2014, it did not simply switch. As the IPUMS CPS documentation records, three eighths of the total sample was randomly selected to receive the redesigned income questions, and the larger portion of the sample, five eighths, was given the existing questions on income.

That produced one year in which both instruments were measured on comparable populations, and it is the only year in which the gap between them is observable.

Two details from that documentation are the practically useful part. The first is the honest statement of the limitation: "Because the 5/8 file and the 3/8 file are not completely comparable, WTSUPP values have been assigned so that either file is individually representative of the entire US population."

The second is the guidance, which is directional. The recommendation is to use the three-eighths file for comparing income estimates from the 2014 supplement with 2015 and beyond, and for anyone comparing 2014 with 2013 and earlier, to use the five-eighths file.

Read that twice, because it is the whole lesson. In the overlap period you do not have one number. You have two, and which one you quote depends on which direction you are comparing. A single headline figure for the transition year is the thing being given up, deliberately, in exchange for both series staying internally valid.

The three honest options

Once you have measured the gap in an overlap wave, there are exactly three defensible things to do, and one indefensible one.

Bridge. Restate the historical series onto the new instrument by applying the measured gap, and label every restated point as adjusted. Appropriate when the gap is well measured and roughly uniform.

Break. Declare the old series closed and start a new one, with a visible discontinuity on the chart and no number spanning the break. Appropriate when the gap is poorly measured or varies by segment. Breaking a series is not a failure; it is the honest reading of what you know.

Dual-run. Keep fielding both versions indefinitely, usually on a reduced sample for the legacy item. Expensive, and worth it only for a small number of genuinely load-bearing metrics tied to targets or compensation.

Splice. Append the new number to the old line and say nothing. This is the option that produces the eight-point improvement that was really one point.

A worked bridge

Run the overlap wave as a random split: half the sample gets the old item, half gets the new one.

  • Q4, old item only: 62
  • Q1 old half: 63. The real movement from Q4 to Q1 is therefore +1.
  • Q1 new half: 70. The instrument gap is 70 - 63 = +7.
  • Q2 onward, new item only: 72, so Q1 to Q2 moved +2 on the new instrument.

Had you spliced Q4 to the Q1 new-item reading, you would have reported 70 - 62 = +8 when the real customer movement was +1. Seven of your eight points were the question.

To bridge, add the gap to each historical point: Q4 restated onto the new instrument is 62 + 7 = 69. The restated series reads 69 then 70 then 72, which tells the true story of a slow climb rather than a step change.

Two reasons your bridge may not survive contact

The gap is an estimate, and usually an imprecise one. A bridge factor is a difference between two proportions, and differences are noisier than levels. At around 65 percent agreement and 400 respondents per arm, the standard error of that difference is about 3.4 points, giving a 95 percent interval of roughly plus or minus 6.6 points on a measured gap of 7. You would be applying a correction that is barely distinguishable from zero. At 1,000 per arm the interval narrows to about plus or minus 4.2 points, and at 2,500 per arm to about plus or minus 2.6.

The consequence is counter-intuitive and worth planning for: the overlap wave needs a bigger sample than a normal wave, because its job is to estimate a difference rather than a level. If you cannot afford that, you cannot afford to bridge, and breaking the series is the correct choice.

The gap may not be one number. Suppose the aggregate gap is +7, but the wording change lands differently by segment: +11 for SMB respondents and +3 for enterprise, with the two groups at equal share. The weighted average is 0.5 times 11 plus 0.5 times 3, which is exactly 7. The aggregate bridge looks perfect while overstating the enterprise series by 4 points and understating the SMB series by 4.

This is the trap worth naming loudly: an aggregate bridge that reconciles beautifully is not evidence the bridge is right. Always compute the gap within your main reporting segments, and if they disagree materially, either bridge per segment or break the series. A reconciling total is exactly what you would see in the bad case.

Running this in Koji

The mechanics of an overlap wave are mostly a sampling and bookkeeping problem, and Koji handles the bookkeeping part.

Stable question identity. Koji questions carry stable IDs that are preserved in templates and are designed to support cross-study comparison. That matters here because it makes an instrument change visible as an event rather than something you reconstruct later from wording memory. If the ID changed, the item changed.

Six constrained question types. Koji offers open_ended, scale, single_choice, multiple_choice, ranking and yes_no, described in the structured questions guide. Each type has a defined report output, a scale item rendering a distribution and a single_choice item a frequency chart. The practical warning is that changing the type is always an instrument change, and usually a larger one than changing the words, as the 58 against 35 open-versus-prompted gap above shows.

Running both arms at once. Field the legacy and revised question sets as two Koji studies in the same window, recruiting from the same source, and treat the pair as one wave. Koji per-interview quality scoring on a 1 to 5 scale, with its relevance, depth and coverage breakdown, is useful for confirming the two arms were comparable in execution before you attribute the gap to wording.

Keep the old item alive deliberately. If a metric is tied to a target, keep the legacy item in Koji on a reduced sample rather than deleting it. The cost of a small dual-run is far lower than the cost of discovering mid-year that your committed number is not comparable to its baseline.

Common mistakes

Changing several things in one release. New wording, new scale labels and a new sample source in one wave produces a gap you cannot attribute. Change one thing per wave.

Treating a small gap as no gap. A gap that is not statistically significant on your overlap sample is an imprecisely measured gap, not a measured zero. Report the interval.

Bridging a scale change arithmetically. Moving from a 5-point to a 7-point scale is not fixed by rescaling. The categories mean different things to respondents; measure the gap empirically or break the series.

Forgetting the annotation. Any restated point must be labelled as adjusted on the chart itself, not in a footnote in an appendix nobody opens.

Assuming respondents did not also change. An instrument gap measured in one wave may not hold years later, because the population and its reference points drift. See adaptation and response shift for why the zero point moves on its own.

Frequently asked questions

What is a bridge study?

A bridge study fields an old and a new version of a question in the same period on randomly assigned subsamples, so that the difference between them can be measured and used to link the historical series to the new one. The term comes from official statistics, where survey redesigns are routinely accompanied by an overlap period for exactly this purpose. Without an overlap period there is no way to separate the effect of the instrument change from real change in the population.

How is this different from a split-ballot experiment?

A split-ballot experiment and a bridge study use the same design: two question versions randomly assigned within one wave. The difference is purpose. A split-ballot is run to learn how sensitive an answer is to its wording. A bridge study is run because you have already decided to change the instrument and you need a linking factor to keep your trend interpretable. The same wave can serve both purposes.

Can I fix a spliced series after the fact?

Generally no. If both versions were never fielded together, the instrument effect and the real change are confounded in a way no later analysis can separate. The honest options are to mark the point at which the question changed as a break in the series, or to run an overlap wave now and bridge forward from this point while leaving the earlier splice flagged as not comparable.

How large should the overlap sample be?

Larger than a normal wave, because you are estimating a difference rather than a level. At roughly 65 percent agreement, 400 respondents per arm gives a 95 percent interval of about plus or minus 6.6 points on the gap, which is too wide to support a seven-point correction. Around 1,000 per arm narrows it to about plus or minus 4.2 points and 2,500 per arm to about plus or minus 2.6. If the required sample is unaffordable, break the series instead.

Should I bridge or break the series?

Bridge when the gap is precisely measured and consistent across your main reporting segments. Break when the gap is imprecise, varies materially by segment, or arises from a scale or question-type change rather than wording. Breaking is the safer default, and a visible discontinuity on a chart costs far less credibility than a trend that later turns out to have been an artefact.

Does changing a question type in Koji count as an instrument change?

Yes, and usually a larger one than rewording. Koji supports six structured question types, and converting an open_ended item into a single_choice item changes what respondents are prompted to consider. Pew Research Center found 58 percent naming the economy when it was offered explicitly against 35 percent volunteering it unprompted. Treat any type change in Koji as requiring its own overlap wave, and note that Koji stable question IDs make the change auditable.

Related Resources

Sources

  • Pew Research Center. Writing Survey Questions. Pew Research Center methods guide.
  • IPUMS CPS. 2014 3/8 ASEC File Information. University of Minnesota.
  • US Census Bureau. Current Population Survey Annual Social and Economic Supplement, 2014 income question redesign.

Related Articles

The Zero Point Moved: Adaptation and Response Shift in Long-Running Research (2026)

Your satisfaction tracker is flat while the product got better. That is not a measurement failure: the internal scale users rate against re-zeroes itself every time you ship.

Conflicting Research Findings: What to Do When Qualitative and Quantitative Data Disagree (2026)

When your interviews say one thing and your analytics say another, averaging them is the worst possible move. A step-by-step protocol for diagnosing and resolving conflicting research findings.

Why Your Quarterly Metric Shows a Trend That Is Not There (2026)

Undersampling does not blur a cycle, it counterfeits a different one. How the gap between your measurement waves manufactures smooth trends, flat lines, and reversed directions - and the three-question test that catches it.

Split-Ballot Experiments: How Much of Your Number Is the Question?

Write two versions of the item, randomly assign half your sample to each, and the gap is the wording effect. The technique that tells you whether your metric is a fact about customers or about your questionnaire.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

How to Write Unbiased Survey Questions: Avoiding Leading, Loaded & Double-Barreled Questions

A practical guide to question wording — the biggest hidden source of bad data. Learn to spot and fix leading, loaded, double-barreled, and assumptive questions, with real research examples and a pre-launch checklist.