Levels of Measurement: Which Statistics Each Survey Question Type Allows (2026)
Nominal, ordinal, interval and ratio data explained for customer research: the summaries each level supports, how Koji's six structured question types map onto them, and a relabelling test that catches meaningless statistics.
Levels of Measurement: Which Statistics Each Survey Question Type Allows (2026)
Answer first: Every answer you collect sits at one of four levels of measurement: nominal, ordinal, interval or ratio. The level decides which summaries mean something and which ones are artefacts of how you happened to number the options. Nominal data (unordered categories) supports counts, percentages and the mode. Ordinal data (ordered categories, which includes most rating scales) adds the median and percentiles. Interval data (equal spacing) adds the mean and standard deviation. Ratio data (equal spacing plus a true zero) adds statements like twice as many. One test replaces the whole taxonomy: a summary is meaningful only if it survives every relabelling of the scale that keeps the information intact. With tools like Koji, each of the six structured question types maps to a level, and the report chooses summaries that match it.
Where the four levels come from
The framework comes from S. S. Stevens, On the Theory of Scales of Measurement, Science 103(2684):677-680, 1946. Stevens' insight was that a scale is defined not by the numbers written on it but by the transformations you could apply to those numbers without losing anything. If any one-to-one relabelling is harmless, the numbers are only names. If only order-preserving relabellings are harmless, the numbers carry rank and nothing more. If only linear rescalings (multiply and add a constant, like Celsius to Fahrenheit) are harmless, the spacing is real but the zero is arbitrary. If only multiplication is harmless (metres to feet), the zero is real too.
| Level | What the numbers encode | Relabelling that changes nothing | Summaries that survive | Research example |
|---|---|---|---|---|
| Nominal | Identity only | Any one-to-one swap of labels | Counts, percentages, mode, chi-square | Plan tier, region, the theme a quote was coded to |
| Ordinal | Order | Any order-preserving change | All of the above plus median, percentiles, top-box share, rank correlations | Satisfaction rating, a ranking, company size bands |
| Interval | Order and equal spacing | Multiply by a positive number and add a constant | All of the above plus mean, standard deviation, Pearson correlation | Temperature in Celsius; rating scales by convention |
| Ratio | Order, equal spacing and a true zero | Multiply by a positive number | All of the above plus ratios, percent change, coefficient of variation | Sessions per week, minutes to complete a task, spend |
Each level inherits every summary from the levels above it in the table. The practical consequence is asymmetric: using a lower-level summary on higher-level data only wastes information, while using a higher-level summary on lower-level data can manufacture a finding.
The one test that replaces the whole table
You do not need to memorise which statistic belongs to which level. You need one habit: before you report a conclusion, relabel the scale in a way that keeps its information, and see whether the conclusion survives.
Here is the most common failure it catches. Segment A averages 4.0 on a 1-to-5 satisfaction scale; segment B averages 2.0. The draft report says A is twice as satisfied as B. Now relabel the same five points as 0 to 4, which carries exactly the same information. A becomes 3.0 and B becomes 1.0, so A is now three times as satisfied. Relabel them 11 to 15 and A is 14.0 against B's 12.0, about 17 percent more. The same data supports twice, three times and 17 percent, depending only on where you started counting. The ratio was never a finding. It was a property of the labels.
The statement A is higher than B survives all three relabellings, because each of them preserves spacing. That is the right conclusion at this level. Whether it survives relabellings that do not preserve spacing is a harder question, and it is the subject of when relabelling the scale reverses which group scores higher.
Mapping Koji's six question types to levels
Koji supports six structured question types, and each produces data at a predictable level. Knowing the mapping at design time tells you what you will be allowed to say in the report.
| Koji question type | Level of the answer | Summaries that are safe | Watch out for |
|---|---|---|---|
open_ended | Nominal, once answers are coded into themes | Theme counts, share of respondents per theme, representative quotes | Averaging theme codes; treating theme order as a ranking |
single_choice | Nominal, or ordinal if the options have a natural order | Share per option; cumulative share when options are ordered | Losing the order of banded options such as company size |
multiple_choice | A set of nominal yes/no answers, one per option | Share of respondents selecting each option | Adding option shares together (they legitimately exceed 100 percent) |
yes_no | Nominal, two categories | Proportion with an interval around it | Reporting a proportion from a tiny base without its interval |
ranking | Ordinal, within each respondent | First-place share, pairwise preferences, median rank | Comparing average rank across lists of different lengths |
scale | Ordinal, analysed as interval by convention | Distribution, median, top-box share, mean with the distribution beside it | Ratios and percent change on the scale points |
Ratio-level data does not come from any rating format. It comes from asking for a quantity: how many times a week someone exports a report, how many minutes a task takes, how much a team spends. If you want to say twice as often, ask for the count. A scale question whose points are literal counts (0 to 10 exports last week) gives you ratio data; a 1-to-5 frequency scale running from never to always does not.
For the full configuration of each type, see the structured questions guide.
What Koji reports for each level
Because the question type is declared up front, Koji's automatic analysis can choose summaries that fit the level instead of applying one template to everything:
- Scale questions report the mean, the median and the full distribution of answers, and on 0-to-10 or 1-to-10 scales Koji also computes the Net Promoter Score automatically. You get the level-appropriate summary (the distribution) next to the convenient one (the mean), so a reader can see when they disagree.
- Single choice and multiple choice questions report the count and percentage for each option.
- Yes/no questions report the split.
- Ranking questions report the average position of each item, and the full per-respondent orderings stay available for the richer summaries described in why adding one option can reverse your ranking results.
- Open-ended answers are grouped into themes with counts and quotes, which turns free text into nominal data you can tabulate.
Every aggregated number links back to the conversations that produced it, so a surprising summary can be checked against what people actually said. That traceability matters more at lower levels of measurement, where a single summary number hides more.
Five level mistakes that show up in real research reports
1. Averaging category codes. If regions are coded 1 to 4 in a spreadsheet, the average region of 2.3 is a number about the coding scheme. Nominal codes support counts and nothing that involves arithmetic.
2. Ratios on rating scales. As the worked example showed, twice as satisfied depends on where the scale starts. Report differences in points, or better, differences in the share of people at the top of the scale.
3. Percent change on a score with an arbitrary zero. Net Promoter Score runs from -100 to +100, and zero means promoters and detractors are equally common, not that loyalty is absent. A move from 10 to 20 is a 10-point gain; calling it NPS doubled is a ratio statement on a scale with no true zero. The NPS survey guide covers the scoring itself.
4. Throwing away order in banded answers. Company size bands, tenure bands and frequency bands are ordinal. Charting them as unordered bars hides the most useful summary, the cumulative share (what fraction is at or below each band).
5. Adding up multiple-choice shares. When respondents can pick several options, the percentages describe separate yes/no decisions and routinely total more than 100. Adding them, or presenting them as a pie, treats a set of nominal variables as one.
Lord's objection: the level is a default, not a law
Stevens' rules drew an early and famous objection. In On the Statistical Treatment of Football Numbers (American Psychologist 8(12):750-751, 1953), Frederic Lord told the story of freshmen who suspected they had been sold jersey numbers lower than the ones sold to sophomores. The numbers were plainly nominal, yet a test on their mean answered the freshmen's question perfectly well, because the question was about the numbers themselves, not about the players.
Lord's point is that the level of measurement is a property of the question you are asking, not a stamp on the data. That is why this article frames the rule as a test (does the conclusion survive a harmless relabelling?) rather than a list of banned statistics. The same test explains why most researchers average rating scales without incident, which is the subject of can you average Likert scale data?, and when that habit genuinely breaks.
Choosing the level at design time
The cheapest time to fix a level problem is before fieldwork, because you cannot recover a level you did not collect:
- Write the sentence you want to put in the report. Enterprise users export twice as often requires ratio data. Enterprise users are more satisfied only requires ordinal data.
- Pick the question type that produces that level. A count for ratio claims, a
scalefor graded judgements,rankingfor forced trade-offs,single_choicefor categories. - Keep order where it exists. If the options have a natural order, say so in the analysis plan and report cumulative shares.
- Pair every structured answer with the reason behind it. A rating tells you where someone sits; it does not tell you why. Koji's AI interviewer asks a follow-up after the structured answer, so a 3 out of 5 arrives with the sentence that explains it. That qualitative context is what lets you interpret an ordinal number without over-reading it.
Traditional form tools such as SurveyMonkey, Typeform or Qualtrics collect the structured answer well. Koji collects the structured answer and the explanation in the same conversation and analyses both automatically, which removes most of the manual coding that a level-aware analysis otherwise needs.
Frequently asked questions
What are the four levels of measurement?
Nominal, ordinal, interval and ratio. Nominal data is unordered categories, ordinal data is ordered categories, interval data has equal spacing between values, and ratio data has equal spacing plus a true zero. Each level supports every summary of the levels below it and adds its own: the median at ordinal, the mean at interval, ratios and percent change at ratio.
Is a Likert or rating scale ordinal or interval?
Strictly, a single rating item is ordinal, because nothing guarantees that the gap between 3 and 4 equals the gap between 4 and 5. In practice most researchers analyse rating scales as interval by convention, and for many comparisons that is defensible. The safe habit is to report the mean alongside the full distribution so that disagreements between them are visible.
Which statistics can I use on nominal data?
Counts, percentages, the mode, and tests of association such as chi-square. Any statistic that involves adding or averaging the codes is meaningless, because the codes are names. Coded open-ended themes, plan tiers and regions are all nominal.
Why is it wrong to say one group is twice as satisfied as another?
Because the ratio changes when you relabel the scale without changing its information. A 4.0 against a 2.0 on a 1-to-5 scale becomes 3.0 against 1.0 on a 0-to-4 scale, so twice becomes three times. Ratio statements need a true zero, which rating scales do not have. Report the difference in points or in top-box share instead.
How do I get ratio-level data in a customer interview?
Ask for a quantity rather than a rating: how many times, how many minutes, how much money. A scale question whose points are literal counts, such as 0 to 10 exports last week, produces ratio data. A frequency scale running from never to always does not, because its labels are ordered categories rather than amounts.
Does Koji choose the right summary for each question type automatically?
Yes. The question type is declared when you design the study, so Koji reports scale questions with mean, median and distribution (plus NPS on 0-to-10 and 1-to-10 scales), choice questions with counts and percentages, ranking questions with average position, and open-ended questions as coded themes with quotes. Every number links back to the conversations behind it.
Related Resources
- Structured Questions Guide - the six question types and how to configure each one
- Can You Average Likert Scale Data? - when the mean of a rating scale is safe and when it is not
- Why Adding One Option Can Reverse Your Ranking Results - analysing ordinal data from ranking questions
- When Relabelling the Scale Reverses Which Group Scores Higher - the dominance check for group comparisons
- Scale Questions Guide - configuring rating scales in Koji
- Likert Scale Research Guide - writing good Likert statements
Related Articles
Why Adding One Option Can Reverse Your Ranking Results (2026)
Average rank, the default summary for ranking questions, depends on which other options are in the list. A worked example of a reversal no respondent caused, and the first-place and pairwise summaries that stay stable.
Can You Average Likert Scale Data? What the Evidence Actually Says (2026)
Yes, in most situations, and the tests will behave. But robustness is about p-values, not meaning: the median can freeze while real change happens, and the mean can rank two groups in the opposite order from every other summary.
Likert Scale Questions: How to Use Rating Scales in User Research
A complete guide to Likert scale questions in user research — what they are, when to use them, how to write them correctly, and how Koji's AI interviews take rating scales further by pairing quantitative scores with qualitative follow-up.
When Relabelling the Scale Reverses Which Group Scores Higher (2026)
Comparing two groups by average rating assumes the scale points are equally spaced. When the groups' answer distributions cross, an equally valid scoring reverses the result. The cumulative dominance check tells you in advance.
Scale Questions in AI Interviews: Measure NPS, CSAT, and Ratings Automatically
Learn how to configure and use scale questions in Koji AI interviews to capture NPS, CSAT, and satisfaction ratings — with automatic probing and aggregated distribution charts in your research report.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.