Is 4.1 Good? How to Build Internal Benchmarks and Percentile Norms
A raw score means nothing on its own. When no industry benchmark fits your metric, build a norm bank from your own history and convert scores to percentile ranks. Here is the method, the arithmetic, and the sample size below which it is noise.
A score is meaningless until it is compared to something. When no published industry benchmark matches your metric, your wording and your sample, the answer is to build a norm bank: a stored distribution of your own past results, against which any new score can be expressed as a percentile rank. Three past waves is not a norm bank. Twelve is a usable one.
"We got 4.1 out of 5. Is that good?" is the most common question in applied research and the one most often answered badly — usually by reaching for whatever industry benchmark can be found within ten minutes, regardless of whether it was produced by anything resembling your method.
Three legitimate ways to answer
| Approach | The comparison | Best for | Fails when |
|---|---|---|---|
| Criterion-referenced | A target you set in advance | Goals, OKRs, launch gates | The target was arbitrary — most are |
| Norm-referenced | A distribution of comparable scores | Judging whether a result is unusual | The distribution was built from a different method |
| Change-referenced | Your own previous measurement | Tracking programmes, before/after | The change is smaller than your detectable effect |
Most teams believe they are doing the second and are in fact doing a broken version of it. Norm-referencing is only valid when the norming distribution was produced by the same instrument, the same wording, the same scale, the same mode and a comparable sample. Change the mode from email to in-product and the comparison quietly stops meaning anything.
Why published industry benchmarks mislead
The published benchmark for your metric is nearly always measured differently from the way you measure it. Six differences do most of the damage:
- Mode effects. Phone, in-product intercept, email and moderated interview produce systematically different scores from identical wording. In-product intercepts, caught mid-task, skew differently from post-hoc email.
- Scale and wording. A 5-point scale is not a 10-point scale rescaled. "Satisfied with" is not "happy with". Response-scale differences alone can move a mean more than a year of product work.
- Sample source. A panel sample, a customer list and an intercept sample are three different populations wearing the same label.
- Question order. What preceded the question shapes it — the reason survey randomization exists.
- Industry mix. A "SaaS" benchmark averaging developer tools and HR software describes neither.
- Self-selection in the benchmark itself. Vendor benchmark reports are built from customers of that vendor who agreed to share data, which is not a random sample of anything.
The rule worth internalising: a benchmark is comparable only if the method matches, not merely the metric name. Two NPS numbers are not comparable because they are both called NPS. Our NPS benchmarks by industry and customer experience benchmarking guides are useful precisely to the extent that you check the method behind the figure before you use it.
What a properly normed instrument looks like
There is a good example of the real thing, and it is instructive because of how much work it took. Sauro and Lewis built a curved grading scale for the System Usability Scale from 241 usability studies. On it, a SUS score of 68 is a C — the 50th percentile — and the top and bottom 15% of the distribution correspond to A and F grades, with finer subdivisions between.
That scale works because SUS is a fixed instrument: ten items, fixed wording, fixed scoring, applied the same way across hundreds of studies. Every condition for valid norm-referencing is satisfied by construction. See the System Usability Scale guide for the instrument itself.
Almost no in-house metric has any of that. Your satisfaction question has been reworded twice, moved from a 5-point to a 7-point scale, and migrated from email to in-product. Which is exactly why you need to build the norm bank yourself — and why the first step is not statistical.
Building a norm bank: seven steps
1. Freeze the instrument. Before anything else, lock the wording, the scale, the anchors and the mode, and write them down. Every future comparison depends on this and every future stakeholder will want to change it. A norm bank with a wording change halfway through is two short norm banks.
2. Log every wave with its metadata. Not just the score. The metadata is what lets you later determine whether two waves are comparable:
| Field | Why you need it |
|---|---|
| Date and wave ID | Ordering, seasonality |
| n (responses) and invitations sent | Precision, and response rate |
| Exact question wording and scale | Detecting silent drift |
| Mode and channel | The largest single source of incomparability |
| Sample source and any quotas | Population definition |
| Segment breakdown | Enables segment norms later |
| Full response distribution | See step 4 |
3. Set a minimum n per wave. A wave that does not meet it goes in the archive, not the norm bank.
4. Store the distribution, not just the mean. This is the step teams skip and later regret. Keep the full frequency count for each scale point. Without it you cannot compute top-box scores retrospectively, cannot recompute if you change your summary statistic, and cannot inspect whether a stable mean is concealing a polarising split.
5. Compute percentile ranks. The standard formula, which handles ties properly:
Percentile rank = (B + 0.5E) / N × 100
where B is the number of past waves scoring below the current score, E is the number scoring exactly equal, and N is the total number of waves in the bank.
Worked example. Twelve stored waves of a 1–5 satisfaction mean: 3.6, 3.7, 3.8, 3.9, 3.9, 4.0, 4.0, 4.1, 4.2, 4.3, 4.4, 4.6. This wave scores 4.1.
- Below 4.1: seven waves (3.6, 3.7, 3.8, 3.9, 3.9, 4.0, 4.0)
- Equal to 4.1: one wave
- PR = (7 + 0.5 × 1) / 12 × 100 = 62.5
So 4.1 sits at roughly the 63rd percentile of your own history — above your median of 4.0, but well inside normal variation. That is a far more honest and more useful answer than "4.1 is good" or "4.1 is below the industry average of 4.3", and you can say it in a sentence a stakeholder understands: this is a slightly better than typical result for us, not an outlier.
6. Re-baseline on a schedule, not opportunistically. Annually is normal. Re-baselining after a bad wave, because the bank now makes the number look worse, is how norm banks lose credibility permanently. Set the schedule in advance and honour it.
7. Publish the rules. The instrument, the minimum n, the percentile formula, and the re-baselining schedule. A norm bank whose rules are not public is a number the loudest stakeholder can argue with.
Top-box, mean, or grand mean?
Your choice of summary statistic changes how much movement you will see, and this catches people out constantly.
Take a 1–5 satisfaction item and move 10% of respondents from a 4 to a 5 — a genuine improvement, and nothing else changes:
| Statistic | Before | After | Movement |
|---|---|---|---|
| Mean | 4.00 | 4.10 | +0.10 |
| Top-box (% giving 5) | 30% | 40% | +10 points |
| Top-two-box (% giving 4 or 5) | 75% | 75% | 0 |
The same real change appears as a rounding error, a dramatic jump, or nothing at all, purely as a function of which number you report. None of the three is wrong; they answer different questions.
| Statistic | Use when | Watch out for |
|---|---|---|
| Mean | You want maximum statistical efficiency and a stable trend | Compresses real movement; assumes interval spacing |
| Top-box | You care about delight and want a sensitive measure | Noisier; needs larger n |
| Top-two-box | You care about "acceptable or better" | Insensitive to shifts inside the box, as above |
Pick one as your headline before you see the data, and store the distribution so you can always compute the others.
The sample size below which a norm bank is noise
A percentile rank computed over waves that are themselves imprecise inherits that imprecision. Two things must be large enough: the n within each wave, and the number of waves.
Within a wave, the relevant quantity is not the margin of error but the smallest difference you could reliably detect between two waves. Comparing two independent waves costs a factor of √2 in standard error, and detecting a difference at 80% power needs about 2.80 standard errors rather than the 1.96 used for a confidence interval — so the minimum detectable effect is roughly 2.02 times the margin of error. A wave of n=400 has a ±4.9pp margin of error and an effective wave-to-wave detectable difference near 9.9pp. This is developed properly in statistical power and minimum detectable effect, and it is the single most useful piece of arithmetic for anyone running a tracking programme.
Practical implications:
- Waves of n < 100 do not belong in a norm bank. Their movement is mostly sampling noise, and including them widens the distribution in a way that makes genuinely unusual results look normal.
- Below about 8–10 waves, quote the range rather than a percentile. With N=5 waves, each wave is worth 20 percentile points; the precision implied by "the 63rd percentile" is fictional.
- Percentile ranks near the middle are stable; those at the extremes are not. The gap between the 45th and 55th percentile is usually trivial. The distance between the 90th and the 98th is often one unusual wave.
See survey sample size for setting n in the first place.
Segment norms beat overall norms
An overall 4.1 can be 4.5 among enterprise customers and 3.4 among SMB, and the overall figure will move whenever your customer mix moves — even when nothing about either group's experience has changed. That is a mix shift masquerading as a finding, and it is one of the most common false alarms in tracking research.
Build norms per segment for the segments you actually make decisions about, and enforce your minimum n within each segment rather than overall. A segment with 30 responses does not get a percentile rank; it gets a range and a caveat. If your mix changes materially between waves, weighting lets you compare like with like.
When you have no history at all
Everyone starts here. Three moves make the first year productive rather than wasted:
- Declare wave one as the baseline explicitly. Not as a judgement, as a reference point. Say so in the report: "this is the baseline; no percentile interpretation is available until we have a distribution."
- Use criterion-referencing in the interim — but derive the criterion rather than inventing it. The most defensible source is qualitative: ask people what would make it a 5.
- Ask why, from the first wave. A norm bank tells you a score is unusual. It never tells you what caused it. If you only start collecting the reasoning once the number moves, you will have a distribution and no explanation.
That third point is where the design of your instrument matters more than the statistics. A tracking survey gives you a number and a shrug. A conversational interview gives you the number and the reason, from the same participant, in the same sitting.
How Koji makes a norm bank practical
The hardest part of norm-referencing is not the arithmetic — it is holding the instrument still for two years while stakeholders ask to reword things, and capturing the explanation alongside the number.
- Structured questions freeze the instrument. All six types —
open_ended,scale,single_choice,multiple_choice,rankingandyes_no— are defined explicitly, and ascalequestion carries a fixed minimum, maximum and labels. Wave twelve asks exactly what wave one asked, and the response distribution is stored as discrete values rather than reconstructed from prose. That is a norm bank you can actually trust. - AI follow-up questions capture the why at the moment the number is given. When a respondent selects a 3, Koji's AI asks what would have made it higher — so every wave in your bank arrives with its own explanation attached. No industry benchmark table can tell you why your score moved; your own participants can, and they will if something asks them.
- Voice and text interviews hold the mode constant across waves while still letting participants choose how to answer — and mode consistency is the single largest threat to comparability in a tracking programme.
- Real-time reports recompute themes, quotes and quantitative summaries as interviews complete, so a wave closes with its distribution ready to store rather than waiting on manual analysis.
The result is that the tedious discipline norm-referencing demands — same instrument, stored distributions, reasoning captured alongside the number — becomes a property of how the study runs rather than a process somebody has to remember.
Frequently asked questions
How many past waves do I need before percentile ranks mean anything? Around 8–10 as a working minimum, and 12 or more for comfortable interpretation. Below that, each wave carries too much percentile weight — with five waves in the bank, one wave moves the answer by 20 points. Report the range and the median instead until the bank is deep enough.
Should I use an industry benchmark or my own history? Your own history, whenever you have enough of it, because it is the only comparison where the method is guaranteed to match. Use industry benchmarks for orientation when entering a new category or setting an initial target, and always check the mode, scale, wording and sample source behind the published figure before you compare against it.
What is the percentile rank formula for a norm bank? Percentile rank = (B + 0.5E) / N × 100, where B is the number of past waves below the current score, E is the number exactly equal, and N is the total number of waves. The half-weighting of ties is what stops identical scores producing different percentiles depending on ordering.
Why did our score drop when nothing changed? Check the mix before you check the experience. If the proportion of responses from a lower-scoring segment increased, the overall score falls without any group's experience changing. This is why segment-level norms matter, and why weighting exists. Also check whether the drop exceeds your minimum detectable effect — roughly twice your margin of error for a wave-over-wave comparison — before treating it as real at all.
Should our headline metric be the mean or top-box? Decide before you see the data, and store the full response distribution either way so you can compute both. Top-box is more sensitive to real change and noisier; the mean is more stable and compresses movement. Moving 10% of respondents from a 4 to a 5 shifts the mean by 0.10, the top-box by 10 points, and the top-two-box by nothing at all.
Can I add old waves to a norm bank if the wording changed? No. A wording, scale or mode change breaks comparability, and including pre-change waves widens your distribution with results that were never measuring the same thing. Start a new bank at the change, note the break in your documentation, and keep the old bank for historical reference only.
Related Resources
- Structured Questions in AI Interviews — the six question types that keep an instrument fixed across waves
- Statistical Power and Minimum Detectable Effect — the arithmetic behind the minimum n for a usable norm bank
- System Usability Scale (SUS) — the best worked example of a properly normed instrument
- NPS Benchmarks by Industry — an industry table, and how to check it before using it
- Customer Experience Benchmarking — measuring against external standards when they genuinely fit
- Brand Tracking Studies — running the wave programme a norm bank is built from
- Survey Weighting: How to Correct a Skewed Sample — comparing like with like when your mix shifts
Want every wave to arrive with its distribution and its explanation? Start free with 10 credits and run a tracking study where the AI asks why the number moved.
Related Articles
Brand Tracking Studies: How to Measure Brand Health Over Time (2026)
A complete guide to brand tracking studies — what to measure, how often to run them, sample size, and how AI-native platforms make continuous brand tracking affordable for the first time.
Customer Experience Benchmarking: How to Measure Against Industry Standards
A complete guide to CX benchmarking — how to measure your customer experience performance against competitors and industry standards using both quantitative metrics and qualitative interviews.
NPS Benchmarks 2026: Net Promoter Score by Industry (Complete Reference)
Compare your NPS to 2026 industry benchmarks for SaaS, ecommerce, financial services, healthcare, and more. Includes what counts as "good", scoring math, and how to dig into the "why" behind your score with AI follow-up interviews.
Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)
Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Survey Weighting: How to Correct a Skewed Sample
A practical guide to survey weighting — post-stratification, raking, and propensity weighting — plus how to calculate design effect and effective sample size, and when weighting cannot save your data.
System Usability Scale (SUS): Complete Guide with Calculator, Benchmarks & Examples
The definitive 2026 guide to the System Usability Scale (SUS): the 10-question formula, scoring calculator, Sauro–Lewis benchmark grades, and how to deploy SUS at scale with AI-moderated interviews on Koji.