{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-20T06:45:14.779Z"},"content":[{"type":"documentation","id":"6f820f91-7a9d-42cf-82cb-1762ebc7d2ae","slug":"ordinal-scale-group-comparison-reversal","title":"When Relabelling the Scale Reverses Which Group Scores Higher (2026)","url":"https://www.koji.so/docs/ordinal-scale-group-comparison-reversal","summary":"Comparing two groups by mean rating assumes equally spaced scale points. When the groups' cumulative distributions cross, some order-preserving rescoring reverses which group is ahead: segment B beats A 3.4 to 3.2 on a default 1-5 scoring, but scoring the top point as 7 puts A ahead 4.0 to 3.4. If one group's cumulative share is at or below the other's at every point (first-order stochastic dominance), the direction is safe under any scoring. When curves cross, report distributions, top-box and bottom-box, and investigate the polarised group.","content":"# When Relabelling the Scale Reverses Which Group Scores Higher (2026)\n\n**Answer first:** When you compare two groups by their average rating, you are silently assuming the points on the scale are equally spaced. If the two groups' answer distributions cross, there is always another spacing, just as consistent with the answers, under which the other group has the higher average. Your data, your arithmetic and your chart can all be correct and the conclusion can still be a property of the labels. The fix is a check you can do in a spreadsheet: **compare cumulative distributions.** If one group sits at or above the other at every point on the scale, the ordering is safe under any spacing. If they cross, no average can settle which group is ahead, and you should report the distributions and the reasons behind them. Koji shows the full distribution for every scale question, so the check is always on the page.\n\n## The worked example\n\nTwo customer segments answer the same 1-to-5 satisfaction question.\n\n| Segment | Answers | Mean | Median | Top-2-box |\n| --- | --- | --- | --- | --- |\n| A | 1, 1, 4, 5, 5 | 3.2 | 4 | 60% |\n| B | 3, 3, 3, 4, 4 | **3.4** | 3 | 40% |\n\nBy the mean, B is more satisfied. Now ask what the numbers 1 to 5 actually claim. They claim the step from *satisfied* to *very satisfied* is the same size as every other step. Suppose instead that very satisfied customers are much further from satisfied ones than the labels imply (they renew, expand and refer), and score the top point as 7 rather than 5. Nothing about the order of the answers changes. The averages become:\n\n| Scoring of the five points | Mean A | Mean B | Who is ahead |\n| --- | --- | --- | --- |\n| 1, 2, 3, 4, 5 (the default) | 3.2 | 3.4 | B |\n| 1, 2, 3, 4, 7 (top point further away) | 4.0 | 3.4 | A |\n| -1, 2, 3, 4, 5 (bottom point further away) | 2.4 | 3.4 | B, by a wider margin |\n\nThree scorings, all order-preserving, all equally faithful to what respondents said. Two point to B, one to A. The median (A ahead) and top-2-box (A ahead) never moved, because they never used the spacing in the first place.\n\nThis is not a contrived edge case. It is the normal shape of a polarised segment against a lukewarm one, and it is the pattern Torrin Liddell and John Kruschke documented at scale in *Analyzing ordinal data with metric models: What could possibly go wrong?* (Journal of Experimental Social Psychology 79:328-348, 2018), including \"systematic inversions of effects, for which treating ordinal data as metric indicates the opposite ordering of means than the true ordering of means\".\n\n## The check that tells you in advance: cumulative dominance\n\nFor each scale point, compute the share of each group at or below that point.\n\n| At or below | Segment A | Segment B | Segment C |\n| --- | --- | --- | --- |\n| 1 | 40% | 0% | 0% |\n| 2 | 40% | 0% | 0% |\n| 3 | 40% | 60% | 40% |\n| 4 | 60% | 100% | 80% |\n| 5 | 100% | 100% | 100% |\n\nA lower cumulative share means more people higher up the scale. Compare A with B: at points 1 and 2, A has more people at the bottom (40 percent against 0); at points 3 and 4, B has more people at or below (60 against 40, then 100 against 60). **The curves cross.** Whenever they cross, some order-preserving scoring puts A ahead and another puts B ahead, which is exactly what the table above showed.\n\nNow compare a third segment, C, answering 3, 3, 4, 4, 5, with B. C's cumulative share is at or below B's at every point (40 against 60 at point 3, 80 against 100 at point 4). **C dominates B.** Under dominance, the direction of the comparison is guaranteed for every order-preserving scoring: C's mean is 3.8 against B's 3.4 on the default scoring, 4.2 against 3.4 with the top point at 7, and 3.8 against 3.4 with the bottom point at -1. The size of the gap changes; the direction cannot.\n\nThis is a standard result from decision theory, where it is called first-order stochastic dominance: one distribution has at least as high an average as another under every increasing scoring if and only if its cumulative curve never rises above the other's. It converts a philosophical argument about levels of measurement into a mechanical check.\n\n## Why nothing inside the analysis warns you\n\nThe unsettling part of the A-versus-B comparison is that no step is wrong. The answers were recorded correctly, the means were computed correctly, and a significance test on them behaves as advertised, which is the robustness Geoff Norman's 2010 review documents (see [can you average Likert scale data?](/docs/can-you-average-likert-scale-data)). The error is not in the data or the arithmetic. It is in the unstated assumption that the gaps between labels are equal, and that assumption is invisible in any output that shows only the mean.\n\nLiddell and Kruschke make the same point: \"there is no sure-fire way to detect these problems by treating the ordinal values as metric\". You have to look at the ordinal structure itself, which is what the cumulative table does. They also report that \"averaging across multiple ordinal measurements does not solve or even ameliorate these problems\", so a multi-item index is not a way out.\n\nThe comparison is a special case of a broader rule from the [levels of measurement](/docs/levels-of-measurement-survey-data) guide: a conclusion is meaningful only if it survives every relabelling that keeps the information intact. For comparisons of rating-scale means, dominance is precisely the condition under which it does.\n\n## What to do when the curves cross\n\nA crossing is not a failed analysis. It is a finding: the two groups differ in shape, not just in level. Report it that way.\n\n1. **Show both distributions side by side** and name the difference: *A is polarised, B is uniformly lukewarm.*\n2. **Report summaries that do not depend on spacing**: top-box, bottom-box and median. Say plainly that they disagree with the mean, and why.\n3. **Split the polarised group.** A crossing usually means one segment contains two populations. The people answering 1 in segment A and the people answering 5 are almost certainly different customers with different needs.\n4. **If a model is required, use one built for ordered outcomes.** Liddell and Kruschke advocate ordered-probit models; ordered-logit is the more common frequentist equivalent. Both estimate where the thresholds between scale points actually sit instead of assuming them.\n5. **Decide which end of the scale the decision depends on.** If the business question is churn, the bottom of the distribution matters most and A is the worry. If it is referral, the top matters and A is the opportunity. The mean averages away the one part you need.\n\n## Getting the reasons behind the shape\n\nThe cumulative table tells you that segment A is split. It cannot tell you why. That requires what the people answering 1 and the people answering 5 actually said.\n\nThis is where an interview beats a form. In Koji, every `scale` answer can be followed by an AI probe, so each band of the distribution arrives with explanations attached, and the report shows the mean, the median and the full answer distribution for every scale question, with each data point linked back to its conversation. A polarised segment shows up in the distribution chart, and the themes behind its 1s and its 5s sit one click away. Koji collects scale answers alongside the other five structured types (`open_ended`, `single_choice`, `multiple_choice`, `ranking` and `yes_no`), so a crossing on a satisfaction question can be cross-referenced against, say, a `single_choice` question on plan tier to find which population is which. See the [structured questions guide](/docs/structured-questions-guide) for how each type is configured.\n\nTraditional survey tools stop at the chart. Platforms like Koji automate the follow-up that turns a crossing from an ambiguous average into two clearly described customer groups.\n\n## A pre-report checklist for any group comparison on a rating scale\n\n- Build the cumulative table for the groups being compared.\n- Dominance? Report the mean difference; the direction is safe under any scoring.\n- Crossing? Report distributions, top-box and bottom-box, and describe the shape difference in words.\n- Check the base sizes; small groups produce crossings from noise alone (see [why the top and bottom segments are the smallest](/docs/segment-ranking-sample-size-artifact)).\n- Pull the reasons from the extremes of any polarised group before presenting a conclusion.\n\n## Frequently asked questions\n\n### Can the average rating of two groups really reverse under a different scoring?\n\nYes, whenever their answer distributions cross. In a worked example, segment B beat segment A on the default 1-to-5 scoring (3.4 against 3.2), but scoring the top point as 7 instead of 5, which keeps every answer in the same order, put A ahead (4.0 against 3.4). Both scorings are equally faithful to what respondents said.\n\n### What is stochastic dominance in survey data?\n\nOne group dominates another when its cumulative share at or below each scale point is never higher than the other group's. In that case the dominant group has the higher or equal average under every order-preserving scoring of the scale, so the direction of the comparison is safe. If the cumulative curves cross, no such guarantee exists.\n\n### How do I check whether my group comparison is safe?\n\nFor each scale point, compute the share of each group answering at or below it, and put the two columns side by side. If one column is at or below the other on every row, the comparison is safe in direction. If the columns swap order at any row, the curves cross and the mean comparison depends on the spacing you assumed.\n\n### Does a significant t-test protect me from this?\n\nNo. A t-test controls how often you find a difference that does not exist; it does not guarantee that the direction of the difference reflects respondents' ordering rather than the spacing of the labels. Liddell and Kruschke note there is no sure-fire way to detect these problems while treating ordinal values as metric.\n\n### What should I report when two groups' distributions cross?\n\nShow both distributions, report top-box, bottom-box and median alongside the mean, and describe the difference in shape in words, for example that one group is polarised and the other uniformly lukewarm. Then investigate the polarised group, which usually contains two different populations of customers.\n\n### How does Koji help with ordinal group comparisons?\n\nKoji reports the full answer distribution next to the mean and median for every scale question, so crossings are visible without extra work. Because the AI interviewer can probe after each rating, the reasons behind the 1s and the 5s of a polarised group are themed in the report, and every number links back to the conversations behind it.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and how to configure each one\n- [Can You Average Likert Scale Data?](/docs/can-you-average-likert-scale-data) - what the robustness evidence does and does not promise\n- [Levels of Measurement](/docs/levels-of-measurement-survey-data) - which statistics each question type allows\n- [Why Adding One Option Can Reverse Your Ranking Results](/docs/average-rank-ranking-question-analysis) - the same trap in ranking data\n- [Measurement Invariance](/docs/measurement-invariance-comparing-groups) - whether two groups read the scale the same way\n- [Extreme Response Bias](/docs/extreme-response-bias) - when polarisation is a response style rather than an opinion\n","category":"Analysis & Synthesis","lastModified":"2026-09-19T03:23:14.446293+00:00","metaTitle":"Comparing Groups on a Rating Scale: The Reversal Check","metaDescription":"Two groups' average ratings can swap order under an equally valid scoring. The cumulative dominance check tells you when a comparison is safe.","keywords":["compare groups likert scale","ordinal data group comparison","stochastic dominance survey","likert mean comparison","ordered probit survey data","rating scale group differences","polarized survey results"],"aiSummary":"Comparing two groups by mean rating assumes equally spaced scale points. When the groups' cumulative distributions cross, some order-preserving rescoring reverses which group is ahead: segment B beats A 3.4 to 3.2 on a default 1-5 scoring, but scoring the top point as 7 puts A ahead 4.0 to 3.4. If one group's cumulative share is at or below the other's at every point (first-order stochastic dominance), the direction is safe under any scoring. When curves cross, report distributions, top-box and bottom-box, and investigate the polarised group.","aiPrerequisites":["Familiarity with rating-scale questions","Experience comparing segments or groups in a report"],"aiLearningOutcomes":["Explain why mean comparisons on rating scales can reverse","Build a cumulative distribution table for two groups","Apply the dominance check before ranking groups","Report crossing distributions as a finding about shape"],"aiDifficulty":"advanced","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}