Mix Shift: Why Your Score Fell When Every Segment Improved (2026)
Your headline metric can fall while every segment inside it improves. Learn the Kitagawa decomposition that splits a metric change into rate and composition components, and how to act on it.
Your headline score can fall in a quarter when every single segment inside it improved. Nothing has to go wrong for this to happen, and nobody has to change their mind. It happens because a headline score is a weighted average, and you changed the weights. Demographers have had the tool for this since 1955: split the change into a rate component (people's scores moved) and a composition component (the mix of people moved). Until you have run that split, you do not know whether your number is telling you about your product or about your go-to-market.
This article shows you how to run the decomposition by hand, how to read it, and — the part most teams skip — how to turn the answer into the one research question worth asking next.
The result that makes the case
Here is a real shape, with numbers you can check. A company measures satisfaction on a 0-10 scale, by plan tier, in two consecutive quarters.
| Plan | Q1 respondents | Q1 score | Q2 respondents | Q2 score |
|---|---|---|---|---|
| Starter | 2,000 | 6.8 | 5,200 | 7.0 |
| Pro | 1,200 | 8.1 | 1,400 | 8.2 |
| Enterprise | 300 | 8.9 | 320 | 9.0 |
| All | 3,500 | 7.43 | 6,920 | 7.34 |
Read the segment rows: Starter is up 0.2, Pro is up 0.1, Enterprise is up 0.1. Every tier improved. Now read the total: 7.43 down to 7.34, a fall of 0.09.
There is no arithmetic error. Starter grew from 57.1% of respondents to 75.1%, and Starter is the lowest-scoring tier. A bigger slice of the total is now drawn from the group that was always least satisfied, and that shift is large enough to swamp three genuine improvements.
This pattern has a formal name — it is an instance of Simpson's paradox — and the most cited demonstration of it is an admissions case. In the autumn of 1973 the Graduate Division at the University of California, Berkeley made admission decisions for 12,763 applicants across 101 departments. The admission rate for the 8,442 male applicants was approximately 44.2%; for the 4,321 female applicants it was approximately 34.6%. Investigating, Bickel, Hammel and O'Connell found that the aggregate gap did not survive disaggregation, because the "proportion of women applicants tends to be high in departments that are hard to get into and low in those that are easy to get into," and concluded there was "no pattern of discrimination on the part of the admissions committee." The gap was in the mix, not in the decisions.
Your quarterly metric is the same object as that admissions rate. It is a single number standing in for a population whose composition you are actively changing every time marketing turns a campaign on.
The decomposition, in the form you can actually compute
Evelyn Kitagawa published the method in "Components of a Difference Between Two Rates" (Journal of the American Statistical Association, 1955, volume 50, pages 1168-1194). It splits the difference between two overall rates into the part attributable to differing composition and the part attributable to differing group-specific rates.
For segments indexed by i, with share w and score r:
- Composition component = sum over i of (w2i - w1i) × the average of r1i and r2i
- Rate component = sum over i of (r2i - r1i) × the average of w1i and w2i
The two components add exactly to the observed change. That is the property that makes the method worth using rather than eyeballing: it is an identity, not an approximation.
Run it on the table above:
| Component | Value | Reading |
|---|---|---|
| Composition | -0.257 | The mix moved toward lower-scoring plans |
| Rate | +0.166 | Scores within plans genuinely improved |
| Total | -0.091 | Matches the observed fall exactly |
Per segment, the composition arithmetic shows where the pressure came from: Starter contributes +1.242 to the composition term, Pro -1.145 and Enterprise -0.353. Starter's own weight rose so much that it dominates, and because the other two tiers lost share, their contributions are negative. Net: -0.257.
So the honest sentence for the board is not "satisfaction declined." It is: "Satisfaction improved in every tier. The reported average fell because self-serve signups grew from 57% to 75% of our respondent base." Those two sentences lead to opposite decisions.
This is not survey weighting
The closest neighbour to this method in most research teams' vocabulary is weighting, and conflating them is the most common way to get this wrong. They solve different problems.
Weighting repairs a sample. You believe your respondents are unrepresentative of a population you know the true shape of, so you re-weight them to match it. The population is the truth and your sample is the error. That is a well-covered discipline — post-stratification, raking, propensity weighting, design effects and weight trimming are all set out in the survey weighting guide.
Decomposition compares two populations. In the table above, nobody is under-represented. Those really are the customers. The 75% Starter share is not a sampling defect to be corrected; it is a fact about the business. There is no "true" mix being approximated, so there is nothing to repair.
The practical test: ask whether you would be upset if you learned the mix was real. If a skewed mix means your fieldwork failed, you have a weighting problem. If a skewed mix means your company changed, you have a decomposition problem. Reaching for weights in the second case silently deletes your actual commercial result.
Where mix shift hides in a product organisation
Composition change is not exotic. It is the normal consequence of a company doing anything at all. The usual sources:
- Acquisition mix. A paid campaign, a pricing change, a free tier, a new geography or a channel partnership. Any of them can move segment shares by tens of points in a single quarter.
- Differential churn. If your least satisfied customers leave fastest, your average satisfaction rises with no product change whatsoever — the survivor version of the same arithmetic, covered in survivorship bias in customer research.
- Differential response. Even with a stable customer base, if one segment starts responding at a higher rate, your respondent mix moves even though your customer mix did not.
- Tenure mix. A growth spurt loads your base with new accounts. If sentiment varies with tenure, your average moves for reasons that are purely about the age structure of the base.
That last one is the doorway to a bigger problem: the mix that shifted may be a mix of time, not of type. That is treated in tenure, calendar, or vintage.
Running the decomposition properly
1. Choose the segmentation before you look. The decomposition is only as meaningful as the partition it runs over, and slicing after seeing the result is how teams manufacture explanations — see the multiple comparisons problem. Pick the segmentation your business actually runs on: plan, region, channel, tenure band.
2. Use the same segmentation in both periods. If the plan names changed, map them explicitly and write the mapping down. A silently redefined segment produces composition effects out of nothing.
3. Check that the components sum. They must equal the observed change exactly. If they do not, you have a bug — usually shares that do not sum to 1, or a segment present in one period only.
4. Report both components, always. A single adjusted number invites the reader to think the mix effect has been dealt with. It has not; it has been moved somewhere else. Reporting -0.257 and +0.166 side by side is more informative than any one number, and it is honest about the fact that two different things happened.
5. Confirm the segments are comparable at all. A decomposition assumes the score means the same thing in each segment. If Starter users interpret a 7 differently from Enterprise buyers, the arithmetic is fine and the interpretation is not — test this with measurement invariance before drawing conclusions from between-segment gaps.
The part the arithmetic cannot do
A decomposition tells you where a change came from. It cannot tell you why, and it cannot tell you whether the change is good.
In the worked example the composition term is -0.257 because self-serve grew. Is that bad? If self-serve is the strategy, a falling headline average is the arithmetic signature of a plan working, and "fix satisfaction" would be exactly the wrong response. If self-serve growth was unintentional — a discount that attracted a poorly fitting audience — the same -0.257 is an early warning. The number is identical in both cases. No amount of further slicing resolves it, because the difference lives in intent and in customer experience, not in the aggregate.
This is where the decomposition should hand off to fieldwork, and it hands off with an unusual advantage: it tells you exactly whom to talk to. You do not need a general satisfaction study. You need the 3,200 additional Starter respondents who did not exist last quarter, and you need to know what they expected when they signed up.
How Koji closes the loop
The classic obstacle is that by the time you have the decomposition, the research to explain it costs more than the analysis did. Recruiting, scheduling and moderating a round of interviews with a newly arrived segment is weeks of work, so most teams write "mix shift" in the appendix and move on.
Platforms like Koji remove that step. Because interviews are AI-moderated and run on the participant's schedule in voice or text, you can put the decomposition's own output straight into fieldwork:
- Target the segment the arithmetic named. Import the exact cohort whose share moved and interview it, rather than fielding a general study and hoping the segment shows up.
- Ask the question the arithmetic cannot answer. Koji's AI asks follow-up questions automatically, so when a new Starter user says the product "wasn't what I expected," the interviewer probes what they did expect — the thing no aggregate contains.
- Keep the quantitative spine. Koji's six structured question types —
open_ended,scale,single_choice,multiple_choice,rankingandyes_no— mean the same study yields both the segment-level scale scores your next decomposition needs and the open-ended explanation for the one you just ran. See the structured questions guide for how to combine them. - Re-run it as a standing check. Because a study can stay open continuously, the composition term becomes a monitored quantity rather than a quarterly surprise.
A traditional survey tool gives you the average that misled you. The point of an AI-native platform is that explaining the average costs days rather than a quarter.
A working checklist
- Recompute the headline metric as a weighted average of segment scores, and confirm it reproduces the reported number.
- Run the Kitagawa split; verify the components sum to the observed change.
- State the result as two sentences: what the rates did, and what the mix did.
- Decide, explicitly, whether the mix change was intended.
- Interview the segment whose share moved — not the whole base.
- Put the composition term on the dashboard next to the headline, permanently.
Frequently asked questions
What is mix shift in a customer metric?
Mix shift is a change in the composition of the population a metric is computed over, as opposed to a change in the behaviour or attitudes of the people in it. Because headline metrics are weighted averages of segment-level values, moving the weights moves the headline even when every segment value is stable or improving. It is the reason an aggregate can fall while every component rises.
How is the Kitagawa decomposition calculated?
Split the total change into two sums across segments. The composition component sums each segment's change in share multiplied by its average score across the two periods. The rate component sums each segment's change in score multiplied by its average share across the two periods. The two components add exactly to the observed change in the overall rate, which gives you a built-in check on the arithmetic. The method comes from Kitagawa's 1955 paper in the Journal of the American Statistical Association.
Is mix shift the same as Simpson's paradox?
They are closely related. Simpson's paradox is the striking case where the aggregate comparison reverses the direction shown in every subgroup. Mix shift is the general mechanism that produces it — a change in group weights driving the aggregate independently of group-specific values. Every Simpson's paradox involves composition change, but plenty of composition change merely dampens or exaggerates a trend without reversing it, which is harder to notice and just as misleading.
Should I just report the mix-adjusted number instead?
No, report both. An adjusted number silently embeds a choice of which mix to treat as the baseline, and different choices produce different answers. Reporting the rate and composition components separately keeps that choice visible. The trap is covered in detail in there is no neutral baseline.
Does this apply to NPS and CSAT specifically?
Yes, and to any rate, ratio or average computed over a heterogeneous population — NPS, CSAT, activation rate, conversion rate, retention, average order value and support ticket rates all behave this way. NPS is particularly exposed because it is already a compressed transformation of a distribution, so a modest change in respondent mix can move it by several points.
How many segments should I decompose over?
Use the segmentation your business already operates on, typically three to six groups, and fix it before looking at results. More segments give a finer attribution but thinner cells and more scope for reading noise as signal. If a segment has too few respondents to have a stable score, merge it rather than reporting a component computed on a handful of people.
Related Resources
- Survey Weighting: How to Correct a Skewed Sample
- There Is No Neutral Baseline: Choosing the Customer Mix You Compare Against
- Tenure, Calendar, or Vintage: The Three Effects Hiding in Every Cohort Chart
- The Multiple Comparisons Problem
- Measurement Invariance: Why You Cannot Compare Scores Across Segments
- Structured Questions Guide
Related Articles
Customer Segmentation Research: How to Build Segments That Actually Drive Decisions
How to use qualitative interviews — rather than demographic surveys — to build behavioral and motivational customer segments that product, marketing, and sales teams actually use.
Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)
Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.
The Multiple Comparisons Problem: Why Slicing Data Into Segments Manufactures Findings (2026)
Test 20 segments at the 5 percent threshold and you have a 64 percent chance of finding at least one difference that is not there. Learn how to count the tests you actually ran, when to control the family-wise error rate versus the false discovery rate, and why a correction cannot rescue a bad prior.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Weighting: How to Correct a Skewed Sample
A practical guide to survey weighting — post-stratification, raking, and propensity weighting — plus how to calculate design effect and effective sample size, and when weighting cannot save your data.
Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story
Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.