Back to docs
Analysis & Synthesis

There Is No Neutral Baseline: Choosing the Customer Mix You Compare Against (2026)

Every mix-adjusted metric embeds a choice of which customer mix counts as standard. Different choices give different answers, and adjusting away a deliberate strategy deletes your result.

Every mix-adjusted number answers the question "what would this metric have been if our customer base had looked like that instead?" — and somebody has to choose what "that" is. There is no neutral choice. Adjust to last quarter's mix and you have privileged the past. Adjust to your target mix and you have privileged the roadmap. Adjust to an even split and you have privileged a customer base you have never had. Each is defensible, each produces a different number, and none of them is the truth.

This is the sequel to a decomposition, and it inverts its advice. Once you know how to separate rate change from composition change, the reflex is to report the composition-adjusted figure and move on. That reflex is wrong often enough to be dangerous — because when the mix change is your strategy working, adjusting it away deletes the result you were trying to measure.

The same data, three defensible answers

Take the worked example from mix shift: satisfaction by plan tier across two quarters, where every tier improved but the headline fell from 7.43 to 7.34.

Now compute the adjusted change — the change you would have seen if the mix had been held fixed — under three different fixed mixes:

Standard mix usedQ1 adjustedQ2 adjustedAdjusted change
Q1 mix (57% / 34% / 9%)7.437.58+0.157
Q2 mix (75% / 20% / 5%)7.167.34+0.175
Equal thirds7.938.07+0.133
No adjustment (crude)7.437.34-0.091

All three adjusted answers agree on the sign, which is reassuring. But they disagree on magnitude by about a third — +0.133 to +0.175 — and the levels disagree far more: Q1 satisfaction is "7.43" or "7.16" or "7.93" depending purely on which mix you declared standard. If your target is a level ("get to 7.5"), the choice of standard decides whether you have hit it.

Nothing in the data selects among these. The standard is an input you supply.

The published case where the ranking moved

If a third of a point sounds tolerable, the epidemiological literature has a much larger demonstration, because public health had this argument at national scale and had to resolve it in public.

Age-standardised mortality rates are computed by applying a country's age-specific death rates to a fixed "standard population." Europe used a standard built in 1976. By 2013 that standard no longer resembled any actual European age structure, so Eurostat published a revision based on projected EU and EFTA populations for 2011-2030.

Tadayon, Wickramasinghe and Townsend examined what the switch did to cardiovascular disease mortality across Europe (Population Health Metrics, 2019, volume 17, article 6). Their findings:

  • "CVD rates calculated using the 1976 ESP were on average half the size of rates calculated using the 2013 ESP (mean rate difference = 1.95; P < 0.001)." The mean rate difference was 1.86 for men and 2.03 for women.
  • Country rankings changed: "ranks of countries by ASMRs calculated using the two ESPs were different for both sexes." The largest single move was males in Kazakhstan, six places lower under the 2013 standard than under the 1976 one.
  • Their conclusion was that "the 2013 ESP changes the relative burden of CVD mortality rates for European countries by sex."

No country's death rates changed. No data changed. A committee changed the reference population, and the league table rearranged itself. That is the cleanest available proof that a standardised number is a joint product of the data and the standard — and it is exactly what happens to your internal segment league tables when someone quietly re-bases the comparison mix.

There is no single world standard either. Segi's world standard population dates from 1960, the WHO published its own world standard in 2001, and Europe maintains the ESP. Rates computed against different standards are not comparable, which is why any competent epidemiological table states its standard in the caption. Almost no product dashboard does.

The sign inversion: when adjusting destroys the finding

Here is the failure that the decomposition habit walks you into.

Suppose the mix shift in the example was deliberate. The company's plan for the year was to move volume into self-serve: cheaper to acquire, faster to activate, a much larger addressable market. Starter grew from 57% to 75% of respondents because the strategy worked.

Now look at what the two reports say:

  • Crude: satisfaction fell 0.09. Reads as a problem.
  • Mix-adjusted: satisfaction rose 0.157. Reads as a success.

Both are misleading, and the adjusted one is worse. The adjusted figure answers: "what would satisfaction have done if we had not executed our strategy?" That is a counterfactual about a company that does not exist. The real result — we deliberately acquired a large volume of customers who are systematically less satisfied than our existing base, and here is what that costs — is invisible in both numbers. It only appears when you report the composition term itself as a finding rather than as a nuisance to be removed.

Norman Ryder, whose 1965 paper established cohort analysis in the social sciences, made precisely this complaint about the habit of standardising. Age, he wrote, is customarily used in statistical analyses "as a cross-sectional nuisance to be controlled by procedures like standardization," and this "implicitly static orientation ignores an important source of variation." The founder of the field warned that the adjustment throws away the signal. (Ryder, "The Cohort as a Concept in the Study of Social Change," American Sociological Review, 1965, volume 30, pages 843-861.)

The rule: adjust away a mix change you did not cause and do not want. Report a mix change you caused on purpose. The arithmetic cannot tell these apart, because intent is not in the data.

How this differs from over-controlling a regression

There is a related trap that this article is deliberately not about. Adding a control variable to a model can create an association that is not there, when the control is a common effect of the exposure and the outcome; that mechanism is collider bias, and it is a distinct failure with a distinct diagnosis.

The problem here is simpler and more specific: standardisation is not producing a spurious association, it is answering a well-defined question that you may not have meant to ask. The number is correct. The question is wrong. Knowing which of the two you are facing matters, because the fix differs — a collider is removed from the model, whereas a standard population is chosen, documented and defended.

Choosing and documenting a standard

Practical guidance that survives contact with a quarterly review:

1. Name the standard in the caption, every time. "CSAT, standardised to the Q1 2026 plan mix" is a complete statement. "CSAT (adjusted)" is not. This one habit prevents most of the damage.

2. Prefer the earlier period as the standard when you are asking "did our product get better?" Holding the mix at its earlier value answers the question a product team is actually asking: for a customer base like the one we had, is the experience better now?

3. Prefer the later period when you are asking "what are we delivering today?" This weights the customers you actually have now, which is the right basis for a resourcing or support-capacity decision.

4. Never let the standard drift silently. Re-basing to the most recent quarter every quarter means each period is compared against a different yardstick, and a multi-quarter trend built that way is not a trend of anything. If you must re-base, restate the whole history.

5. Report the crude number alongside. The crude number is what your customers, your support queue and your revenue actually experienced. It is not a worse number; it is a different question. Suppressing it in favour of an adjusted figure is how a company convinces itself that a real deterioration in delivered experience is a composition artefact.

6. When the mix change is strategic, make it the headline. "We grew self-serve from 57% to 75% of the base; that segment runs 1.2 points below Pro on satisfaction; the blended figure therefore fell 0.09 while every tier improved" is one sentence and it is the whole story.

The question a standard population cannot settle

Every option above is a way of reporting a mix change. None of them tells you whether the newly arrived customers are less satisfied because they are worse fits, because they were promised something different, or because the product genuinely serves them badly. Those three have identical arithmetic signatures and completely different remedies.

That question is not answerable by choosing a better baseline. It is answerable by asking the people in the segment that grew.

The modern approach: making the comparison group affordable

The reason teams reach for a single adjusted number is usually economics rather than conviction. Explaining a composition change means fielding research against a specific, often newly arrived, segment — and against the segment it displaced, so you have something to compare with. Two targeted studies at traditional cost and lead time is a quarter of a research team's capacity, so the adjusted number goes in the deck instead.

Platforms like Koji change that arithmetic. AI-moderated interviews run in voice or text without a moderator present, so fielding the comparison group costs roughly what fielding the first group costs:

  • Interview both mixes, not one. Run the same guide against the segment that grew and the segment that shrank. The difference between them is the thing your standard population was trying to approximate — except now it is measured rather than assumed.
  • Ask about expectations, not just satisfaction. Koji's AI asks follow-up questions automatically, so a Starter user who rates 7 gets asked what would have made it a 9, and what they thought they were buying.
  • Keep the segment structure in the instrument. Using single_choice and multiple_choice questions to capture plan, channel and use case, scale for the metric itself, ranking for priorities, yes_no for eligibility screens, and open_ended for the explanation means every study returns data already shaped for the next decomposition. The structured questions guide covers combining them.
  • Standing studies instead of quarterly snapshots. When a study can stay open, the composition question gets answered continuously rather than being reconstructed after the fact from a mix you can no longer interview.

The honest summary of this whole topic is that a standardised metric replaces a hard question with a tractable one. That was a reasonable trade when talking to a few hundred customers took six weeks. It is a poor trade when it takes two days.

A working checklist

  • Write down which mix you standardised to, in the chart caption.
  • Recompute under at least one alternative standard and check whether the sign, not just the magnitude, is stable.
  • Ask whether the composition change was intended. If yes, report it as a result.
  • Never re-base silently; restate history if you re-base at all.
  • Publish the crude number next to the adjusted one.
  • Interview the segment that grew and the one it displaced.

Frequently asked questions

What is a standard population?

A standard population is a fixed set of group weights used to compute a comparable summary rate across populations or time periods. Each group's rate is multiplied by its weight in the standard rather than its weight in the actual population, so differences in composition cannot drive the comparison. In demography and epidemiology the standard is usually an age distribution; in product research it is usually a plan, segment, channel or tenure distribution.

Which standard population should I use for product metrics?

Use the mix from the earlier period when the question is whether the experience improved for a comparable customer base, and the mix from the later period when the question is what you are delivering today. Whichever you pick, state it explicitly and keep it fixed across the whole time series. The worst option is re-basing every period, because that compares each quarter against a different yardstick.

Does the choice of standard population really change conclusions?

Yes, and it has done so at national scale. When Europe replaced its 1976 standard population with the 2013 revision, age-standardised cardiovascular mortality rates roughly doubled on average, and the ranking of countries changed — males in Kazakhstan moved six places. No underlying death rate changed. The same mechanism operates on any internal segment league table when the comparison mix is changed.

Is adjusting for mix the same as controlling for a confounder?

Not necessarily, and this is where teams get into trouble. Adjustment is appropriate when the composition difference is a nuisance you did not choose. It is inappropriate when the composition change is itself the effect you are trying to measure — for example when a deliberate go-to-market shift moved your customer mix. In that case adjusting removes the result. Intent is not visible in the data, so this judgement cannot be automated.

Should I ever report only the adjusted number?

Rarely. The crude number is what your customers, support team and revenue actually experienced, so it is a legitimate answer to a legitimate question. Reporting only the adjusted figure invites the reader to believe the composition effect has been dealt with, when it has only been made invisible. Report both, with the standard named.

How do I explain this to stakeholders without a statistics lecture?

Use one sentence with both components in it: "every tier improved, and the blended score fell, because the lowest-scoring tier grew from 57% to 75% of our customers." Stakeholders do not need the decomposition algebra; they need to know that two things happened at once and that one of them was the plan. If they then ask whether the growth was worth it, that is a research question, not an analysis question.

Related Resources

Related Articles

Collider Bias: When Adding a Control Variable Creates the Correlation (2026)

Most research advice tells you to control for more variables. Collider bias is the case where controlling, filtering or segmenting manufactures an association that does not exist. Here is how to recognise it before it reaches a roadmap.

Is 4.1 Good? How to Build Internal Benchmarks and Percentile Norms

A raw score means nothing on its own. When no industry benchmark fits your metric, build a norm bank from your own history and convert scores to percentile ranks. Here is the method, the arithmetic, and the sample size below which it is noise.

Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)

Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.

NPS Benchmarks 2026: Net Promoter Score by Industry (Complete Reference)

Compare your NPS to 2026 industry benchmarks for SaaS, ecommerce, financial services, healthcare, and more. Includes what counts as "good", scoring math, and how to dig into the "why" behind your score with AI follow-up interviews.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Weighting: How to Correct a Skewed Sample

A practical guide to survey weighting — post-stratification, raking, and propensity weighting — plus how to calculate design effect and effective sample size, and when weighting cannot save your data.