Back to docs
Analysis & Synthesis

Can You Average Likert Scale Data? What the Evidence Actually Says (2026)

Yes, in most situations, and the tests will behave. But robustness is about p-values, not meaning: the median can freeze while real change happens, and the mean can rank two groups in the opposite order from every other summary.

Can You Average Likert Scale Data? What the Evidence Actually Says (2026)

Answer first: Yes, in most situations you can average Likert and rating-scale data, and the tests you run on those averages will behave. Decades of simulation work summarised by Geoff Norman (2010) show that parametric statistics hold up well on rating data. But that robustness is about p-values, not about meaning, and it leaves two real risks. The median can stay frozen while real movement happens underneath it, and when two groups' answer distributions cross, the mean can point in the opposite direction from the median and from the top-box share. The working rule: average freely, but never report a rating-scale mean without its distribution, and check for crossing before you rank groups. Koji reports the mean, the median and the full distribution for every scale question by default, so the check is already on the page.

The two camps

The argument over averaging Likert data is old enough to have settled positions.

The strict position follows from the levels of measurement framework. A single rating item is ordinal: strongly agree is above agree, but nothing guarantees the step from agree to strongly agree is the same size as the step from neutral to agree. Means and standard deviations assume equal steps, so on this view they are not permitted. Susan Jamieson's short, much-cited Likert scales: how to (ab)use them (Medical Education 38(12):1217-1218, 2004) is the standard reference for this camp, which recommends medians, modes and non-parametric tests.

The pragmatic position is set out in Geoff Norman's Likert scales, levels of measurement and the "laws" of statistics (Advances in Health Sciences Education 15(5):625-632, 2010). Norman reviews evidence dating back to the 1930s and concludes that "parametric statistics are robust with respect to violations of these assumptions", including small samples, non-normal distributions and ordinal response formats. On this view, refusing to average rating data throws away power for no gain.

In practice the pragmatic camp won. Torrin Liddell and John Kruschke surveyed three leading psychology journals and found that "100% of the articles that analyzed ordinal data did so using a metric model" (Analyzing ordinal data with metric models: What could possibly go wrong?, Journal of Experimental Social Psychology 79:328-348, 2018). Almost everyone averages. The useful question is not whether it is allowed but what it hides.

What robustness does and does not promise

Norman's evidence answers a specific question: if there is truly no difference between two groups, does a t-test on rating data falsely report one more often than it should? Mostly, no. That is a statement about the error rates of a test.

It is not a statement that the mean describes the data well, or that the ordering of two means reflects the ordering of what respondents actually feel. Those are different properties, and they can fail while the test behaves perfectly. Two failures matter in everyday research.

Failure one: the median is too coarse to see change

On a 5-point scale the median can only take a handful of values, so it is often frozen while real movement happens underneath it. Take ten respondents answering a satisfaction question:

WaveAnswersMeanMedianShare answering 5
Before2, 3, 3, 4, 4, 4, 4, 4, 5, 53.8420%
After2, 3, 3, 4, 4, 4, 5, 5, 5, 54.0440%

Two people moved from 4 to 5. The share of fully satisfied customers doubled, from 20 to 40 percent. The mean rose by 0.2. The median did not move at all. A team that followed the strict camp and reported only medians would conclude nothing happened.

This is the strongest practical argument for averaging: on short scales the median discards exactly the movement that product and CX teams care about. It is also the argument for reporting a top-box share, which captured the change most clearly of the three.

Failure two: the mean can disagree with everything else

The second failure runs the other way. When two groups have differently shaped answer distributions, the mean can rank them in the opposite order from the median and the top-box share:

GroupAnswersMeanMedianTop-2-box (4 or 5)
A1, 1, 4, 5, 53.2460%
B3, 3, 3, 4, 43.4340%

By mean, B is ahead. By median and by top-2-box, A is ahead. Neither summary is miscalculated. Group A is polarised, with some people who love the product and some who are very unhappy, while B is uniformly lukewarm. Which group is more satisfied is not a question the numbers can settle on their own, because the answer depends on how far apart you believe the scale points are. Liddell and Kruschke document the general version of this: treating ordinal data as metric can produce inversions, in which the analysis "indicates the opposite ordering of means than the true ordering of means".

The full mechanics, and a simple check that tells you in advance whether a comparison is vulnerable, are in when relabelling the scale reverses which group scores higher.

Does averaging several items fix it?

A common defence is that a multi-item scale (the sum or average of, say, five Likert items) behaves more like interval data than any single item. That is partly true for reliability: summing items reduces noise, which is why validated instruments use several. But Liddell and Kruschke tested the specific claim and report that "averaging across multiple ordinal measurements does not solve or even ameliorate these problems". Averaging several ordinal items produces a finer-grained number, not an interval one.

Which summary to report, and when

SummaryBest atBlind toReport it when
MeanDetecting small shifts; feeding tests and modelsPolarisation; unequal spacing between pointsTracking one group over time, with the distribution beside it
MedianResisting outliersMovement within a scale point; most change on short scalesScales of 7 or more points, or skewed data
Top-box or top-2-box sharePlain-language reporting; sensitivity at the top of the scaleMovement lower down the scaleStakeholder reporting, targets, before-and-after comparisons
Full distributionEverything above, plus shapeNothing, but it is harder to read at a glanceAlways, as the reference the other three are checked against

A sensible default for any scale question is therefore: the distribution as the primary exhibit, top-box share as the headline, the mean as the tracking number, and the median when the scale is long or skewed. Koji's report puts the distribution, the mean and the median on the page for every scale question, so choosing between them is a reading decision rather than an extra analysis step.

A decision rule you can apply in two minutes

  1. Averaging one group over time? Use the mean, and show the distribution for the first and last wave.
  2. Comparing two groups? Compare the cumulative distributions first. If one group is at or above the other at every scale point, any summary will agree on the direction, and the mean is safe to use. If they cross, report the distributions and say what the crossing means. In Koji each group's distribution is already charted, so the comparison takes a glance.
  3. Reporting to non-researchers? Lead with top-box share, which does not depend on the spacing between points.
  4. Running a formal test? A t-test on the means is usually fine for error rates, which is Norman's point. For modelling ordinal outcomes where the direction of an effect matters, an ordered-probit or ordered-logit model describes the data better.

Where the reasons come from

Every summary above tells you what moved. None tells you why. In a form tool, a 3 out of 5 arrives as a bare number and you guess the reason. In Koji, the AI interviewer follows each scale answer with a probe (what would have made that a 4?), so the distribution in the report comes with the explanations behind each band. For a polarised result like group A above, that is the difference between reporting an ambiguous average and knowing which customers are unhappy and why.

Koji's report shows the mean, the median and the full answer distribution for every scale question, and computes the Net Promoter Score automatically on 0-to-10 and 1-to-10 scales, so the level-appropriate summary and the convenient one sit side by side. Scale questions sit alongside the other five structured types (open_ended, single_choice, multiple_choice, ranking and yes_no), all described in the structured questions guide. Because an AI moderator runs every conversation, adding the follow-up costs no extra researcher time, which is what makes the reasons practical to collect at survey scale.

Frequently asked questions

Can you calculate a mean for Likert scale data?

Yes. Strictly a single Likert item is ordinal, but decades of research summarised by Norman (2010) show parametric statistics are robust on rating data, and almost every published study averages them. The caveat is that the mean should never be reported alone: show the distribution beside it, because the mean hides polarisation and can disagree with the median and top-box share.

Should I report the mean or the median for a 5-point scale?

Usually the mean plus the distribution, with a top-box share for headlines. On a 5-point scale the median can only take a few values, so it often stays frozen while real change happens: in a worked example, two of ten respondents moving from 4 to 5 doubled the top-box share and raised the mean by 0.2 while the median did not move.

Is it wrong to run a t-test on Likert data?

Not for most purposes. Simulation evidence going back decades shows t-tests keep close to their nominal error rates on rating data. What a t-test cannot tell you is whether the ordering of two means reflects the ordering of what people feel, which fails when the two groups' answer distributions cross.

Does combining several Likert items into a scale make it interval data?

It makes the score finer-grained and more reliable, but not interval. Liddell and Kruschke (2018) report that averaging across multiple ordinal measurements does not solve or even ameliorate the inversion and error problems they document. Use multi-item scales for reliability, not as a licence to ignore distribution shape.

What is a top-box score and why use it?

Top-box is the share of respondents choosing the highest scale point; top-2-box includes the second highest too. It does not depend on the spacing between scale points, it is easy for non-researchers to read, and it is often more sensitive than the median to change at the top of the scale.

How does Koji report Likert and scale questions?

Every scale question in a Koji report shows the mean, the median and the full distribution of answers, and 0-to-10 or 1-to-10 scales also get an automatic Net Promoter Score. Because the AI interviewer probes after each rating, the report also carries the reasons behind each part of the distribution, with every number linked back to the conversations it came from.

Related Resources

Related Articles

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.

Ceiling and Floor Effects: When Your Scale Cannot Measure the Change You Care About (2026)

If more than 15 percent of respondents score the maximum, your metric has gone blind - and it goes blind first on your best customers. Learn how to run a headroom audit, why ceilings manufacture false segment differences, and which question types have no ceiling at all.

Levels of Measurement: Which Statistics Each Survey Question Type Allows (2026)

Nominal, ordinal, interval and ratio data explained for customer research: the summaries each level supports, how Koji's six structured question types map onto them, and a relabelling test that catches meaningless statistics.

Likert Scale Questions: How to Use Rating Scales in User Research

A complete guide to Likert scale questions in user research — what they are, when to use them, how to write them correctly, and how Koji's AI interviews take rating scales further by pairing quantitative scores with qualitative follow-up.

When Relabelling the Scale Reverses Which Group Scores Higher (2026)

Comparing two groups by average rating assumes the scale points are equally spaced. When the groups' answer distributions cross, an equally valid scoring reverses the result. The cumulative dominance check tells you in advance.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.