When Relabelling the Scale Reverses Which Group Scores Higher (2026)
Comparing two groups by average rating assumes the scale points are equally spaced. When the groups' answer distributions cross, an equally valid scoring reverses the result. The cumulative dominance check tells you in advance.
When Relabelling the Scale Reverses Which Group Scores Higher (2026)
Answer first: When you compare two groups by their average rating, you are silently assuming the points on the scale are equally spaced. If the two groups' answer distributions cross, there is always another spacing, just as consistent with the answers, under which the other group has the higher average. Your data, your arithmetic and your chart can all be correct and the conclusion can still be a property of the labels. The fix is a check you can do in a spreadsheet: compare cumulative distributions. If one group sits at or above the other at every point on the scale, the ordering is safe under any spacing. If they cross, no average can settle which group is ahead, and you should report the distributions and the reasons behind them. Koji shows the full distribution for every scale question, so the check is always on the page.
The worked example
Two customer segments answer the same 1-to-5 satisfaction question.
| Segment | Answers | Mean | Median | Top-2-box |
|---|---|---|---|---|
| A | 1, 1, 4, 5, 5 | 3.2 | 4 | 60% |
| B | 3, 3, 3, 4, 4 | 3.4 | 3 | 40% |
By the mean, B is more satisfied. Now ask what the numbers 1 to 5 actually claim. They claim the step from satisfied to very satisfied is the same size as every other step. Suppose instead that very satisfied customers are much further from satisfied ones than the labels imply (they renew, expand and refer), and score the top point as 7 rather than 5. Nothing about the order of the answers changes. The averages become:
| Scoring of the five points | Mean A | Mean B | Who is ahead |
|---|---|---|---|
| 1, 2, 3, 4, 5 (the default) | 3.2 | 3.4 | B |
| 1, 2, 3, 4, 7 (top point further away) | 4.0 | 3.4 | A |
| -1, 2, 3, 4, 5 (bottom point further away) | 2.4 | 3.4 | B, by a wider margin |
Three scorings, all order-preserving, all equally faithful to what respondents said. Two point to B, one to A. The median (A ahead) and top-2-box (A ahead) never moved, because they never used the spacing in the first place.
This is not a contrived edge case. It is the normal shape of a polarised segment against a lukewarm one, and it is the pattern Torrin Liddell and John Kruschke documented at scale in Analyzing ordinal data with metric models: What could possibly go wrong? (Journal of Experimental Social Psychology 79:328-348, 2018), including "systematic inversions of effects, for which treating ordinal data as metric indicates the opposite ordering of means than the true ordering of means".
The check that tells you in advance: cumulative dominance
For each scale point, compute the share of each group at or below that point.
| At or below | Segment A | Segment B | Segment C |
|---|---|---|---|
| 1 | 40% | 0% | 0% |
| 2 | 40% | 0% | 0% |
| 3 | 40% | 60% | 40% |
| 4 | 60% | 100% | 80% |
| 5 | 100% | 100% | 100% |
A lower cumulative share means more people higher up the scale. Compare A with B: at points 1 and 2, A has more people at the bottom (40 percent against 0); at points 3 and 4, B has more people at or below (60 against 40, then 100 against 60). The curves cross. Whenever they cross, some order-preserving scoring puts A ahead and another puts B ahead, which is exactly what the table above showed.
Now compare a third segment, C, answering 3, 3, 4, 4, 5, with B. C's cumulative share is at or below B's at every point (40 against 60 at point 3, 80 against 100 at point 4). C dominates B. Under dominance, the direction of the comparison is guaranteed for every order-preserving scoring: C's mean is 3.8 against B's 3.4 on the default scoring, 4.2 against 3.4 with the top point at 7, and 3.8 against 3.4 with the bottom point at -1. The size of the gap changes; the direction cannot.
This is a standard result from decision theory, where it is called first-order stochastic dominance: one distribution has at least as high an average as another under every increasing scoring if and only if its cumulative curve never rises above the other's. It converts a philosophical argument about levels of measurement into a mechanical check.
Why nothing inside the analysis warns you
The unsettling part of the A-versus-B comparison is that no step is wrong. The answers were recorded correctly, the means were computed correctly, and a significance test on them behaves as advertised, which is the robustness Geoff Norman's 2010 review documents (see can you average Likert scale data?). The error is not in the data or the arithmetic. It is in the unstated assumption that the gaps between labels are equal, and that assumption is invisible in any output that shows only the mean.
Liddell and Kruschke make the same point: "there is no sure-fire way to detect these problems by treating the ordinal values as metric". You have to look at the ordinal structure itself, which is what the cumulative table does. They also report that "averaging across multiple ordinal measurements does not solve or even ameliorate these problems", so a multi-item index is not a way out.
The comparison is a special case of a broader rule from the levels of measurement guide: a conclusion is meaningful only if it survives every relabelling that keeps the information intact. For comparisons of rating-scale means, dominance is precisely the condition under which it does.
What to do when the curves cross
A crossing is not a failed analysis. It is a finding: the two groups differ in shape, not just in level. Report it that way.
- Show both distributions side by side and name the difference: A is polarised, B is uniformly lukewarm.
- Report summaries that do not depend on spacing: top-box, bottom-box and median. Say plainly that they disagree with the mean, and why.
- Split the polarised group. A crossing usually means one segment contains two populations. The people answering 1 in segment A and the people answering 5 are almost certainly different customers with different needs.
- If a model is required, use one built for ordered outcomes. Liddell and Kruschke advocate ordered-probit models; ordered-logit is the more common frequentist equivalent. Both estimate where the thresholds between scale points actually sit instead of assuming them.
- Decide which end of the scale the decision depends on. If the business question is churn, the bottom of the distribution matters most and A is the worry. If it is referral, the top matters and A is the opportunity. The mean averages away the one part you need.
Getting the reasons behind the shape
The cumulative table tells you that segment A is split. It cannot tell you why. That requires what the people answering 1 and the people answering 5 actually said.
This is where an interview beats a form. In Koji, every scale answer can be followed by an AI probe, so each band of the distribution arrives with explanations attached, and the report shows the mean, the median and the full answer distribution for every scale question, with each data point linked back to its conversation. A polarised segment shows up in the distribution chart, and the themes behind its 1s and its 5s sit one click away. Koji collects scale answers alongside the other five structured types (open_ended, single_choice, multiple_choice, ranking and yes_no), so a crossing on a satisfaction question can be cross-referenced against, say, a single_choice question on plan tier to find which population is which. See the structured questions guide for how each type is configured.
Traditional survey tools stop at the chart. Platforms like Koji automate the follow-up that turns a crossing from an ambiguous average into two clearly described customer groups.
A pre-report checklist for any group comparison on a rating scale
- Build the cumulative table for the groups being compared.
- Dominance? Report the mean difference; the direction is safe under any scoring.
- Crossing? Report distributions, top-box and bottom-box, and describe the shape difference in words.
- Check the base sizes; small groups produce crossings from noise alone (see why the top and bottom segments are the smallest).
- Pull the reasons from the extremes of any polarised group before presenting a conclusion.
Frequently asked questions
Can the average rating of two groups really reverse under a different scoring?
Yes, whenever their answer distributions cross. In a worked example, segment B beat segment A on the default 1-to-5 scoring (3.4 against 3.2), but scoring the top point as 7 instead of 5, which keeps every answer in the same order, put A ahead (4.0 against 3.4). Both scorings are equally faithful to what respondents said.
What is stochastic dominance in survey data?
One group dominates another when its cumulative share at or below each scale point is never higher than the other group's. In that case the dominant group has the higher or equal average under every order-preserving scoring of the scale, so the direction of the comparison is safe. If the cumulative curves cross, no such guarantee exists.
How do I check whether my group comparison is safe?
For each scale point, compute the share of each group answering at or below it, and put the two columns side by side. If one column is at or below the other on every row, the comparison is safe in direction. If the columns swap order at any row, the curves cross and the mean comparison depends on the spacing you assumed.
Does a significant t-test protect me from this?
No. A t-test controls how often you find a difference that does not exist; it does not guarantee that the direction of the difference reflects respondents' ordering rather than the spacing of the labels. Liddell and Kruschke note there is no sure-fire way to detect these problems while treating ordinal values as metric.
What should I report when two groups' distributions cross?
Show both distributions, report top-box, bottom-box and median alongside the mean, and describe the difference in shape in words, for example that one group is polarised and the other uniformly lukewarm. Then investigate the polarised group, which usually contains two different populations of customers.
How does Koji help with ordinal group comparisons?
Koji reports the full answer distribution next to the mean and median for every scale question, so crossings are visible without extra work. Because the AI interviewer can probe after each rating, the reasons behind the 1s and the 5s of a polarised group are themed in the report, and every number links back to the conversations behind it.
Related Resources
- Structured Questions Guide - the six question types and how to configure each one
- Can You Average Likert Scale Data? - what the robustness evidence does and does not promise
- Levels of Measurement - which statistics each question type allows
- Why Adding One Option Can Reverse Your Ranking Results - the same trap in ranking data
- Measurement Invariance - whether two groups read the scale the same way
- Extreme Response Bias - when polarisation is a response style rather than an opinion
Related Articles
Why Adding One Option Can Reverse Your Ranking Results (2026)
Average rank, the default summary for ranking questions, depends on which other options are in the list. A worked example of a reversal no respondent caused, and the first-place and pairwise summaries that stay stable.
Can You Average Likert Scale Data? What the Evidence Actually Says (2026)
Yes, in most situations, and the tests will behave. But robustness is about p-values, not meaning: the median can freeze while real change happens, and the mean can rank two groups in the opposite order from every other summary.
Extreme Response Bias: Why Some Respondents Always Pick the Extremes
Extreme response bias (ERS) is the tendency of some respondents to over-use the endpoints of a rating scale regardless of the question. Learn why it happens, how much it distorts your data, and how to design scales and AI follow-ups that capture real opinion.
Levels of Measurement: Which Statistics Each Survey Question Type Allows (2026)
Nominal, ordinal, interval and ratio data explained for customer research: the summaries each level supports, how Koji's six structured question types map onto them, and a relabelling test that catches meaningless statistics.
Measurement Invariance: Why You Cannot Compare Scores Across Segments, Languages, or Time (Until You Test This) (2026)
Every segment leaderboard, country comparison and quarterly trend line assumes your questions mean the same thing to everyone. Measurement invariance is the test of that assumption - and it usually fails. Here is what breaks, and what to do about it.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.