{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-20T06:57:51.723Z"},"content":[{"type":"documentation","id":"114191b5-a457-4402-8f81-e823ce702d5e","slug":"can-you-average-likert-scale-data","title":"Can You Average Likert Scale Data? What the Evidence Actually Says (2026)","url":"https://www.koji.so/docs/can-you-average-likert-scale-data","summary":"You can usually average Likert data: Norman (2010) summarises evidence back to the 1930s that parametric statistics are robust on rating data, and Liddell and Kruschke (2018) found 100 percent of surveyed psychology articles did so. Robustness concerns error rates, not meaning. On short scales the median can freeze while the top-box share doubles, and when two groups' distributions cross the mean can rank them opposite to the median and top-box share. Report the mean with the full distribution, and check for crossing before comparing groups.","content":"# Can You Average Likert Scale Data? What the Evidence Actually Says (2026)\n\n**Answer first:** Yes, in most situations you can average Likert and rating-scale data, and the tests you run on those averages will behave. Decades of simulation work summarised by Geoff Norman (2010) show that parametric statistics hold up well on rating data. But that robustness is about p-values, not about meaning, and it leaves two real risks. The median can stay frozen while real movement happens underneath it, and when two groups' answer distributions cross, the mean can point in the opposite direction from the median and from the top-box share. The working rule: **average freely, but never report a rating-scale mean without its distribution, and check for crossing before you rank groups.** Koji reports the mean, the median and the full distribution for every scale question by default, so the check is already on the page.\n\n## The two camps\n\nThe argument over averaging Likert data is old enough to have settled positions.\n\n**The strict position** follows from the [levels of measurement](/docs/levels-of-measurement-survey-data) framework. A single rating item is ordinal: *strongly agree* is above *agree*, but nothing guarantees the step from *agree* to *strongly agree* is the same size as the step from *neutral* to *agree*. Means and standard deviations assume equal steps, so on this view they are not permitted. Susan Jamieson's short, much-cited *Likert scales: how to (ab)use them* (Medical Education 38(12):1217-1218, 2004) is the standard reference for this camp, which recommends medians, modes and non-parametric tests.\n\n**The pragmatic position** is set out in Geoff Norman's *Likert scales, levels of measurement and the \"laws\" of statistics* (Advances in Health Sciences Education 15(5):625-632, 2010). Norman reviews evidence dating back to the 1930s and concludes that \"parametric statistics are robust with respect to violations of these assumptions\", including small samples, non-normal distributions and ordinal response formats. On this view, refusing to average rating data throws away power for no gain.\n\nIn practice the pragmatic camp won. Torrin Liddell and John Kruschke surveyed three leading psychology journals and found that \"100% of the articles that analyzed ordinal data did so using a metric model\" (*Analyzing ordinal data with metric models: What could possibly go wrong?*, Journal of Experimental Social Psychology 79:328-348, 2018). Almost everyone averages. The useful question is not whether it is allowed but what it hides.\n\n## What robustness does and does not promise\n\nNorman's evidence answers a specific question: if there is truly no difference between two groups, does a t-test on rating data falsely report one more often than it should? Mostly, no. That is a statement about the error rates of a test.\n\nIt is not a statement that the mean describes the data well, or that the ordering of two means reflects the ordering of what respondents actually feel. Those are different properties, and they can fail while the test behaves perfectly. Two failures matter in everyday research.\n\n## Failure one: the median is too coarse to see change\n\nOn a 5-point scale the median can only take a handful of values, so it is often frozen while real movement happens underneath it. Take ten respondents answering a satisfaction question:\n\n| Wave | Answers | Mean | Median | Share answering 5 |\n| --- | --- | --- | --- | --- |\n| Before | 2, 3, 3, 4, 4, 4, 4, 4, 5, 5 | 3.8 | 4 | 20% |\n| After | 2, 3, 3, 4, 4, 4, 5, 5, 5, 5 | 4.0 | 4 | 40% |\n\nTwo people moved from 4 to 5. The share of fully satisfied customers doubled, from 20 to 40 percent. The mean rose by 0.2. The median did not move at all. A team that followed the strict camp and reported only medians would conclude nothing happened.\n\nThis is the strongest practical argument for averaging: on short scales the median discards exactly the movement that product and CX teams care about. It is also the argument for reporting a top-box share, which captured the change most clearly of the three.\n\n## Failure two: the mean can disagree with everything else\n\nThe second failure runs the other way. When two groups have differently shaped answer distributions, the mean can rank them in the opposite order from the median and the top-box share:\n\n| Group | Answers | Mean | Median | Top-2-box (4 or 5) |\n| --- | --- | --- | --- | --- |\n| A | 1, 1, 4, 5, 5 | 3.2 | 4 | 60% |\n| B | 3, 3, 3, 4, 4 | 3.4 | 3 | 40% |\n\nBy mean, B is ahead. By median and by top-2-box, A is ahead. Neither summary is miscalculated. Group A is polarised, with some people who love the product and some who are very unhappy, while B is uniformly lukewarm. Which group is *more satisfied* is not a question the numbers can settle on their own, because the answer depends on how far apart you believe the scale points are. Liddell and Kruschke document the general version of this: treating ordinal data as metric can produce inversions, in which the analysis \"indicates the opposite ordering of means than the true ordering of means\".\n\nThe full mechanics, and a simple check that tells you in advance whether a comparison is vulnerable, are in [when relabelling the scale reverses which group scores higher](/docs/ordinal-scale-group-comparison-reversal).\n\n## Does averaging several items fix it?\n\nA common defence is that a multi-item scale (the sum or average of, say, five Likert items) behaves more like interval data than any single item. That is partly true for reliability: summing items reduces noise, which is why validated instruments use several. But Liddell and Kruschke tested the specific claim and report that \"averaging across multiple ordinal measurements does not solve or even ameliorate these problems\". Averaging several ordinal items produces a finer-grained number, not an interval one.\n\n## Which summary to report, and when\n\n| Summary | Best at | Blind to | Report it when |\n| --- | --- | --- | --- |\n| Mean | Detecting small shifts; feeding tests and models | Polarisation; unequal spacing between points | Tracking one group over time, with the distribution beside it |\n| Median | Resisting outliers | Movement within a scale point; most change on short scales | Scales of 7 or more points, or skewed data |\n| Top-box or top-2-box share | Plain-language reporting; sensitivity at the top of the scale | Movement lower down the scale | Stakeholder reporting, targets, before-and-after comparisons |\n| Full distribution | Everything above, plus shape | Nothing, but it is harder to read at a glance | Always, as the reference the other three are checked against |\n\nA sensible default for any scale question is therefore: the distribution as the primary exhibit, top-box share as the headline, the mean as the tracking number, and the median when the scale is long or skewed. Koji's report puts the distribution, the mean and the median on the page for every scale question, so choosing between them is a reading decision rather than an extra analysis step.\n\n## A decision rule you can apply in two minutes\n\n1. **Averaging one group over time?** Use the mean, and show the distribution for the first and last wave.\n2. **Comparing two groups?** Compare the cumulative distributions first. If one group is at or above the other at every scale point, any summary will agree on the direction, and the mean is safe to use. If they cross, report the distributions and say what the crossing means. In Koji each group's distribution is already charted, so the comparison takes a glance.\n3. **Reporting to non-researchers?** Lead with top-box share, which does not depend on the spacing between points.\n4. **Running a formal test?** A t-test on the means is usually fine for error rates, which is Norman's point. For modelling ordinal outcomes where the direction of an effect matters, an ordered-probit or ordered-logit model describes the data better.\n\n## Where the reasons come from\n\nEvery summary above tells you *what* moved. None tells you *why*. In a form tool, a 3 out of 5 arrives as a bare number and you guess the reason. In Koji, the AI interviewer follows each `scale` answer with a probe (*what would have made that a 4?*), so the distribution in the report comes with the explanations behind each band. For a polarised result like group A above, that is the difference between reporting an ambiguous average and knowing which customers are unhappy and why.\n\nKoji's report shows the mean, the median and the full answer distribution for every scale question, and computes the Net Promoter Score automatically on 0-to-10 and 1-to-10 scales, so the level-appropriate summary and the convenient one sit side by side. Scale questions sit alongside the other five structured types (`open_ended`, `single_choice`, `multiple_choice`, `ranking` and `yes_no`), all described in the [structured questions guide](/docs/structured-questions-guide). Because an AI moderator runs every conversation, adding the follow-up costs no extra researcher time, which is what makes the reasons practical to collect at survey scale.\n\n## Frequently asked questions\n\n### Can you calculate a mean for Likert scale data?\n\nYes. Strictly a single Likert item is ordinal, but decades of research summarised by Norman (2010) show parametric statistics are robust on rating data, and almost every published study averages them. The caveat is that the mean should never be reported alone: show the distribution beside it, because the mean hides polarisation and can disagree with the median and top-box share.\n\n### Should I report the mean or the median for a 5-point scale?\n\nUsually the mean plus the distribution, with a top-box share for headlines. On a 5-point scale the median can only take a few values, so it often stays frozen while real change happens: in a worked example, two of ten respondents moving from 4 to 5 doubled the top-box share and raised the mean by 0.2 while the median did not move.\n\n### Is it wrong to run a t-test on Likert data?\n\nNot for most purposes. Simulation evidence going back decades shows t-tests keep close to their nominal error rates on rating data. What a t-test cannot tell you is whether the ordering of two means reflects the ordering of what people feel, which fails when the two groups' answer distributions cross.\n\n### Does combining several Likert items into a scale make it interval data?\n\nIt makes the score finer-grained and more reliable, but not interval. Liddell and Kruschke (2018) report that averaging across multiple ordinal measurements does not solve or even ameliorate the inversion and error problems they document. Use multi-item scales for reliability, not as a licence to ignore distribution shape.\n\n### What is a top-box score and why use it?\n\nTop-box is the share of respondents choosing the highest scale point; top-2-box includes the second highest too. It does not depend on the spacing between scale points, it is easy for non-researchers to read, and it is often more sensitive than the median to change at the top of the scale.\n\n### How does Koji report Likert and scale questions?\n\nEvery scale question in a Koji report shows the mean, the median and the full distribution of answers, and 0-to-10 or 1-to-10 scales also get an automatic Net Promoter Score. Because the AI interviewer probes after each rating, the report also carries the reasons behind each part of the distribution, with every number linked back to the conversations it came from.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) - the six question types and how to configure each one\n- [Levels of Measurement](/docs/levels-of-measurement-survey-data) - which statistics each question type allows\n- [When Relabelling the Scale Reverses Which Group Scores Higher](/docs/ordinal-scale-group-comparison-reversal) - the dominance check for group comparisons\n- [Likert Scale Research Guide](/docs/likert-scale-research-guide) - writing good Likert statements\n- [5-Point vs 7-Point Likert Scale](/docs/5-point-vs-7-point-likert-scale) - choosing the number of scale points\n- [Ceiling and Floor Effects](/docs/ceiling-floor-effects-research) - when a scale runs out of room\n- [Response timing and hedged answers](/docs/dispreferred-responses-interview-timing) - the same trap of reading a difference of means between overlapping distributions as a rule.\n","category":"Analysis & Synthesis","lastModified":"2026-09-20T03:26:06.996759+00:00","metaTitle":"Can You Average Likert Scale Data? Mean vs Median","metaDescription":"Yes, usually. What the evidence on averaging Likert data shows, when the mean misleads, and which summary to report for rating-scale questions.","keywords":["can you average likert scale data","likert scale mean or median","likert data analysis","is likert ordinal or interval","t-test on likert data","top box score","rating scale analysis"],"aiSummary":"You can usually average Likert data: Norman (2010) summarises evidence back to the 1930s that parametric statistics are robust on rating data, and Liddell and Kruschke (2018) found 100 percent of surveyed psychology articles did so. Robustness concerns error rates, not meaning. On short scales the median can freeze while the top-box share doubles, and when two groups' distributions cross the mean can rank them opposite to the median and top-box share. Report the mean with the full distribution, and check for crossing before comparing groups.","aiPrerequisites":["Familiarity with Likert or rating-scale questions","Basic understanding of mean and median"],"aiLearningOutcomes":["Decide when averaging rating-scale data is appropriate","Explain what robustness evidence does and does not guarantee","Choose between mean, median, top-box and distribution for a report","Recognise when a mean comparison between groups is unreliable"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}