Back to docs
Research Methods

Trust the Delta, Not the Level: Common-Mode Bias in Research (2026)

A bias that hits every arm equally cancels when you subtract and survives intact in any absolute number. How to tell which kind of bias you have.

Short answer: a bias that lands on every arm of your study equally largely cancels when you subtract one arm from the other, and survives completely when you report a single absolute number. That is why the same dataset can support a trustworthy "the new flow is 9 points better" and an untrustworthy "our satisfaction is 72%". Instrumentation engineering has a name for the distinction - common-mode rejection - and a hard-won warning attached to it: you have to prove the bias is shared before you are allowed to assume it cancels.

The practical rule: comparisons you ran yourself, under identical conditions, are far more robust than any absolute level you publish. Not because comparisons are magic, but because subtraction removes whatever both sides had in common.

What common-mode rejection is

An electrical engineer measuring a small signal in a noisy room does not measure one wire against ground. They measure two wires against each other, because "the electrical noise from the environment appears as an offset on both input leads, making it a common-mode voltage signal" - and a differential measurement subtracts it away.

The ability to do this is quantified. Common-mode rejection ratio is "a metric used to quantify the ability of the device to reject common-mode signals, i.e. those that appear simultaneously and in-phase on both inputs", defined as CMRR = 20 log10(Ad / |Acm|) dB. Good instrumentation is extraordinarily effective at it: single-chip instrumentation amplifiers "achieve a CMRR in excess of 100 dB, sometimes even 130 dB". A CMRR of 100 dB means the shared interference is attenuated by a factor of 100,000 before it reaches your reading.

The engineering lesson is not "noise is harmless". It is that the same noise has radically different consequences depending on whether you ask an absolute question or a differential one.

The import, stated honestly

Research measurement has the same two question types and rarely distinguishes them.

An absolute question is "what is our task success rate?", "what is our NPS?", "how satisfied are our users?". The answer is a level, and every bias in your instrument is baked into it at full strength.

A differential question is "did the redesign improve success?", "is segment A more frustrated than segment B?", "did this move since last quarter?". The answer is a difference, and any bias that applied equally to both sides has been subtracted out.

This is not an argument that comparisons are safe. It is a much narrower and more useful claim: bias in an absolute number and bias in a difference are different quantities, and you should stop treating a single "is this data any good?" verdict as answering both. Your data can be entirely adequate for the comparison you ran and entirely inadequate for the benchmark number in the same report.

It also explains a pattern every researcher has seen: internally consistent tracking studies that tell a believable story about direction, attached to headline numbers nobody can reconcile with anyone else's.

The rule: subtract and the shared part goes away

Survey mode is the best-documented common-mode candidate in research, because it is a property of the instrument rather than of the respondent.

Pew Research Center ran the comparison properly in 2015, asking 60 identical questions by telephone and on the web. The finding: "differences in responses by survey mode are fairly common, but typically not large, with a mean difference of 5.5 percentage points and a median difference of five points across the 60 questions."

Five and a half points is enough to destroy an absolute claim and usually not enough to flip a direction. If your instrument shifts every reading by about five points, then "72% satisfied" is worthless as a cross-company fact and "up 9 points since March, same instrument" is still informative. Pew also names the mechanism, which is what makes it predictable rather than random: "Respondents may feel a need to present themselves in a more positive light to an interviewer, leading to an overstatement of socially desirable behaviors and attitudes and an understatement of opinions and behaviors they fear would elicit disapproval from another person."

The catch: most real biases are not purely common-mode

Here is where the analogy earns its keep instead of flattering you. A real amplifier has a finite CMRR, and a real survey bias is rarely perfectly shared. When the bias lands unequally on your arms, it does not cancel - and it corrupts the difference, which is the number you were relying on.

Pew's own data contains a clean example. Some questions moved far more than the 5.5-point average: items asking respondents "to assess the quality of their family and social life produced differences of 18 and 14 percentage points, respectively". A bias that is 5 points on one question and 18 on another is not a constant offset. If you compare a family-life question against a social-life question across modes, the mode effect does not subtract out, because the two sides did not receive the same dose.

The Clinton case: how to tell the difference in your own data

The sharpest instance in the Pew study is worth doing the arithmetic on, because the arithmetic is the test.

Overall, "19% of respondents told interviewers they have a 'very unfavorable' opinion of Clinton; that number jumps to 27% on the Web" - a shift of 8 points. If that were a pure common-mode offset, every subgroup would move by about 8 points and all the gaps between subgroups would survive intact.

They did not. Among Republicans and Republican leaners, "fully 53%" held a very unfavorable view on the web "compared with only 36% on the phone" - a shift of 17 points, roughly double the overall shift. So the gap between Republicans and the general population widened from 17 points on the phone to 26 points on the web. The mode changed the comparison, not just the level.

That gives you a concrete diagnostic you can run on any dataset:

  1. Compute the shift for the whole sample.
  2. Compute the shift for each major subgroup.
  3. If the subgroup shifts are similar, the bias is behaving as common-mode, and your subgroup comparisons are defensible even though your levels are not.
  4. If the subgroup shifts differ materially, the bias is differential, and neither your levels nor your comparisons survive without adjustment.

Most teams never run steps 2 to 4 and therefore never learn which regime they are in, which is one reason Koji reports segment-level results by default rather than on request. Note the shape of this failure, which recurs throughout research measurement: checking your absolute numbers against each other cannot detect a differential bias, because the check and the error live on different axes. It is the same structural blind spot described in No Summary Preserves Everything.

A decision rule for absolute numbers

Publish an absolute level only when at least one of these holds:

  • The number is a count of something externally verifiable - tickets filed, tasks completed, refunds issued - rather than a self-report.
  • You are comparing against your own prior measurement on an unchanged instrument, in which case you should report the delta and treat the level as a bookkeeping detail. NN/g's Kate Moran frames the good version of this plainly: "In 2019, the average time to make a purchase was 58 seconds. After our recent redesign, the average time to make a purchase is now 43 seconds."
  • You have run the subgroup-shift diagnostic above and shown the bias is behaving as common-mode.

Be especially wary of the cross-company form, "our success rate for application completion is 86%, while our competitor's is 62%". A 24-point gap sounds decisive, and it is only meaningful if both numbers came off the same instrument under the same conditions - which, for a competitor's number, they almost never did.

How Koji handles this

Common-mode rejection in electronics depends on the two inputs being treated identically. The research equivalent is instrument consistency, and that is precisely where human-moderated research leaks: different moderators, different days, different phrasings, different amounts of rapport. Every one of those differences converts a harmless shared bias into a differential one that will not cancel.

  • An AI moderator holds the instrument constant. Koji asks the same questions in the same way for every participant and every wave. Whatever bias the instrument carries, it carries equally - which is the condition that makes a comparison trustworthy. This is the single most underrated methodological benefit of AI-moderated interviews.
  • Structured questions make the differential test possible. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - asked of everyone. Because a scale or yes_no question is asked identically of every segment, you can actually compute the per-subgroup shifts in steps 2 to 4 above. Emergent, whoever-mentioned-it data cannot support that diagnostic at all.
  • Re-running a wave is cheap, so "same instrument, measured twice" stops being a luxury. Deltas on an unchanged instrument are the most defensible number in research, and Koji makes them the easy number to produce.
  • Automatic thematic analysis and real-time reporting apply one consistent analysis pass rather than a different analyst per wave, removing another source of differential drift.

Legacy survey tools happily hand you an absolute score with a decimal point and no way to check whether the instrument moved between waves. An AI-native platform can hold the instrument still, which is worth more than the decimal point. You do not need a PhD in psychometrics to benefit from it - you need the same questions, asked the same way, twice.

Frequently asked questions

What is common-mode bias in research?

It is a bias that affects every arm of your study by the same amount - every condition, every segment, every wave. Because it is shared, it largely disappears when you subtract one arm from another, and it remains at full strength in any single absolute number you report. Survey mode and question framing are common candidates.

Does this mean A/B tests are immune to bias?

No. It means they are immune to the shared component of bias. Anything that hit both arms equally subtracts out; anything that hit them unequally does not, and randomisation is what makes the equal case likely rather than guaranteed. A bias introduced after assignment, such as one arm being tested by a different moderator, is differential and will corrupt the comparison.

Can I compare my score to an industry benchmark?

Only if the benchmark was collected on an instrument comparable to yours, which is rare. Pew found a mean mode difference of 5.5 percentage points on identical questions, so a benchmark gathered by phone is not interchangeable with your web number. Treat external benchmarks as orientation and your own historical deltas as evidence.

How do I test whether a bias is common-mode?

Compute the shift for your whole sample, then for each major subgroup, and compare. Similar shifts indicate common-mode behaviour and your comparisons survive. Materially different shifts indicate a differential bias that corrupts comparisons too. Koji makes this practical because structured questions are asked of every segment, so per-subgroup shifts are directly computable.

Is this the same as a control group?

They are closely related - a control group is the mechanism by which you create a differential measurement - but the concept is broader. Common-mode thinking also tells you which reported numbers are safe once the study is done, including which levels you should refuse to publish even from a well-controlled study.

Which metrics are safest to report as absolute numbers?

Externally verifiable counts of behaviour: completions, errors, refunds, tickets. Self-reported levels such as satisfaction, likelihood-to-recommend and perceived ease are the least safe, because they are the most sensitive to mode and framing. Report those as deltas on an unchanged instrument wherever you can.

Related Resources

Related Articles

Brand Tracking Studies: How to Measure Brand Health Over Time (2026)

A complete guide to brand tracking studies — what to measure, how often to run them, sample size, and how AI-native platforms make continuous brand tracking affordable for the first time.

The Framing Effect in Surveys and Research: How Question Wording Reverses Answers

The framing effect means the same question, worded as a gain or a loss, produces opposite answers. Learn how framing distorts surveys and interviews — and how neutral, AI-moderated question design keeps your data honest.

When Relabelling the Scale Reverses Which Group Scores Higher (2026)

Comparing two groups by average rating assumes the scale points are equally spaced. When the groups' answer distributions cross, an equally valid scoring reverses the result. The cumulative dominance check tells you in advance.

Statistical Significance in Survey Research: A Plain-English Guide (2026)

A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Nothing Was Measured Twice: Why the Only Error Bar You Can Compute Is the Smallest One

Your margin of error covers sampling and nothing else, because sampling is the only step of a typical study that gets repeated. The Type A and Type B distinction, the definitional floor, and the cheapest ways to buy back replication.