{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-22T17:54:10.094Z"},"content":[{"type":"documentation","id":"4dbd1a49-9ac8-4b65-98c6-a895fcacf741","slug":"internal-benchmarks-percentile-norms","title":"Is 4.1 Good? How to Build Internal Benchmarks and Percentile Norms","url":"https://www.koji.so/docs/internal-benchmarks-percentile-norms","summary":"When no published industry benchmark matches your instrument, build a norm bank: a stored distribution of your own past waves against which new scores become percentile ranks. Published benchmarks mislead because of mode effects, scale and wording differences, sample source, question order, industry mix and self-selection in the benchmark itself - a benchmark is comparable only if the method matches, not the metric name. The Sauro-Lewis SUS curved grading scale, built from 241 studies where 68 equals a C at the 50th percentile, works because the instrument is fixed. Seven steps: freeze the instrument, log every wave with metadata, set a minimum n, store the full distribution not just the mean, compute percentile rank as (B + 0.5E)/N x 100, re-baseline on a schedule, publish the rules. Waves under n=100 do not belong in the bank and fewer than 8-10 waves should be reported as a range; minimum detectable effect is roughly 2.02 times margin of error for wave-over-wave comparison. Moving 10% of respondents from 4 to 5 shifts the mean 0.10, top-box 10 points and top-two-box zero.","content":"**A score is meaningless until it is compared to something. When no published industry benchmark matches your metric, your wording and your sample, the answer is to build a norm bank: a stored distribution of your own past results, against which any new score can be expressed as a percentile rank. Three past waves is not a norm bank. Twelve is a usable one.**\n\n\"We got 4.1 out of 5. Is that good?\" is the most common question in applied research and the one most often answered badly — usually by reaching for whatever industry benchmark can be found within ten minutes, regardless of whether it was produced by anything resembling your method.\n\n## Three legitimate ways to answer\n\n| Approach | The comparison | Best for | Fails when |\n|---|---|---|---|\n| **Criterion-referenced** | A target you set in advance | Goals, OKRs, launch gates | The target was arbitrary — most are |\n| **Norm-referenced** | A distribution of comparable scores | Judging whether a result is unusual | The distribution was built from a different method |\n| **Change-referenced** | Your own previous measurement | Tracking programmes, before/after | The change is smaller than your detectable effect |\n\nMost teams believe they are doing the second and are in fact doing a broken version of it. Norm-referencing is only valid when the norming distribution was produced by *the same instrument, the same wording, the same scale, the same mode and a comparable sample*. Change the mode from email to in-product and the comparison quietly stops meaning anything.\n\n## Why published industry benchmarks mislead\n\nThe published benchmark for your metric is nearly always measured differently from the way you measure it. Six differences do most of the damage:\n\n1. **Mode effects.** Phone, in-product intercept, email and moderated interview produce systematically different scores from identical wording. In-product intercepts, caught mid-task, skew differently from post-hoc email.\n2. **Scale and wording.** A 5-point scale is not a 10-point scale rescaled. \"Satisfied with\" is not \"happy with\". Response-scale differences alone can move a mean more than a year of product work.\n3. **Sample source.** A panel sample, a customer list and an intercept sample are three different populations wearing the same label.\n4. **Question order.** What preceded the question shapes it — the reason [survey randomization](/docs/survey-randomization-guide) exists.\n5. **Industry mix.** A \"SaaS\" benchmark averaging developer tools and HR software describes neither.\n6. **Self-selection in the benchmark itself.** Vendor benchmark reports are built from customers of that vendor who agreed to share data, which is not a random sample of anything.\n\nThe rule worth internalising: **a benchmark is comparable only if the method matches, not merely the metric name.** Two NPS numbers are not comparable because they are both called NPS. Our [NPS benchmarks by industry](/docs/nps-benchmarks-by-industry-2026) and [customer experience benchmarking](/docs/customer-experience-benchmarking) guides are useful precisely to the extent that you check the method behind the figure before you use it.\n\n## What a properly normed instrument looks like\n\nThere is a good example of the real thing, and it is instructive because of how much work it took. Sauro and Lewis built a **curved grading scale for the System Usability Scale from 241 usability studies**. On it, a SUS score of **68 is a C — the 50th percentile** — and the top and bottom 15% of the distribution correspond to A and F grades, with finer subdivisions between.\n\nThat scale works because SUS is a fixed instrument: ten items, fixed wording, fixed scoring, applied the same way across hundreds of studies. Every condition for valid norm-referencing is satisfied by construction. See the [System Usability Scale guide](/docs/system-usability-scale-guide) for the instrument itself.\n\nAlmost no in-house metric has any of that. Your satisfaction question has been reworded twice, moved from a 5-point to a 7-point scale, and migrated from email to in-product. Which is exactly why you need to build the norm bank yourself — and why the first step is not statistical.\n\n## Building a norm bank: seven steps\n\n**1. Freeze the instrument.** Before anything else, lock the wording, the scale, the anchors and the mode, and write them down. Every future comparison depends on this and every future stakeholder will want to change it. A norm bank with a wording change halfway through is two short norm banks.\n\n**2. Log every wave with its metadata.** Not just the score. The metadata is what lets you later determine whether two waves are comparable:\n\n| Field | Why you need it |\n|---|---|\n| Date and wave ID | Ordering, seasonality |\n| n (responses) and invitations sent | Precision, and response rate |\n| Exact question wording and scale | Detecting silent drift |\n| Mode and channel | The largest single source of incomparability |\n| Sample source and any quotas | Population definition |\n| Segment breakdown | Enables segment norms later |\n| Full response distribution | See step 4 |\n\n**3. Set a minimum n per wave.** A wave that does not meet it goes in the archive, not the norm bank.\n\n**4. Store the distribution, not just the mean.** This is the step teams skip and later regret. Keep the full frequency count for each scale point. Without it you cannot compute top-box scores retrospectively, cannot recompute if you change your summary statistic, and cannot inspect whether a stable mean is concealing a polarising split.\n\n**5. Compute percentile ranks.** The standard formula, which handles ties properly:\n\n**Percentile rank = (B + 0.5E) / N × 100**\n\nwhere B is the number of past waves scoring below the current score, E is the number scoring exactly equal, and N is the total number of waves in the bank.\n\nWorked example. Twelve stored waves of a 1–5 satisfaction mean: 3.6, 3.7, 3.8, 3.9, 3.9, 4.0, 4.0, 4.1, 4.2, 4.3, 4.4, 4.6. This wave scores 4.1.\n\n- Below 4.1: seven waves (3.6, 3.7, 3.8, 3.9, 3.9, 4.0, 4.0)\n- Equal to 4.1: one wave\n- PR = (7 + 0.5 × 1) / 12 × 100 = **62.5**\n\nSo 4.1 sits at roughly the 63rd percentile of your own history — above your median of 4.0, but well inside normal variation. That is a far more honest and more useful answer than \"4.1 is good\" or \"4.1 is below the industry average of 4.3\", and you can say it in a sentence a stakeholder understands: *this is a slightly better than typical result for us, not an outlier.*\n\n**6. Re-baseline on a schedule, not opportunistically.** Annually is normal. Re-baselining after a bad wave, because the bank now makes the number look worse, is how norm banks lose credibility permanently. Set the schedule in advance and honour it.\n\n**7. Publish the rules.** The instrument, the minimum n, the percentile formula, and the re-baselining schedule. A norm bank whose rules are not public is a number the loudest stakeholder can argue with.\n\n## Top-box, mean, or grand mean?\n\nYour choice of summary statistic changes how much movement you will see, and this catches people out constantly.\n\nTake a 1–5 satisfaction item and move **10% of respondents from a 4 to a 5** — a genuine improvement, and nothing else changes:\n\n| Statistic | Before | After | Movement |\n|---|---|---|---|\n| Mean | 4.00 | 4.10 | +0.10 |\n| Top-box (% giving 5) | 30% | 40% | **+10 points** |\n| Top-two-box (% giving 4 or 5) | 75% | 75% | **0** |\n\nThe same real change appears as a rounding error, a dramatic jump, or nothing at all, purely as a function of which number you report. None of the three is wrong; they answer different questions.\n\n| Statistic | Use when | Watch out for |\n|---|---|---|\n| Mean | You want maximum statistical efficiency and a stable trend | Compresses real movement; assumes interval spacing |\n| Top-box | You care about delight and want a sensitive measure | Noisier; needs larger n |\n| Top-two-box | You care about \"acceptable or better\" | Insensitive to shifts inside the box, as above |\n\nPick one as your headline before you see the data, and store the distribution so you can always compute the others.\n\n## The sample size below which a norm bank is noise\n\nA percentile rank computed over waves that are themselves imprecise inherits that imprecision. Two things must be large enough: the **n within each wave**, and the **number of waves**.\n\nWithin a wave, the relevant quantity is not the margin of error but the smallest difference you could reliably detect between two waves. Comparing two independent waves costs a factor of √2 in standard error, and detecting a difference at 80% power needs about 2.80 standard errors rather than the 1.96 used for a confidence interval — so **the minimum detectable effect is roughly 2.02 times the margin of error**. A wave of n=400 has a ±4.9pp margin of error and an effective wave-to-wave detectable difference near 9.9pp. This is developed properly in [statistical power and minimum detectable effect](/docs/statistical-power-minimum-detectable-effect), and it is the single most useful piece of arithmetic for anyone running a tracking programme.\n\nPractical implications:\n\n- **Waves of n < 100 do not belong in a norm bank.** Their movement is mostly sampling noise, and including them widens the distribution in a way that makes genuinely unusual results look normal.\n- **Below about 8–10 waves, quote the range rather than a percentile.** With N=5 waves, each wave is worth 20 percentile points; the precision implied by \"the 63rd percentile\" is fictional.\n- **Percentile ranks near the middle are stable; those at the extremes are not.** The gap between the 45th and 55th percentile is usually trivial. The distance between the 90th and the 98th is often one unusual wave.\n\nSee [survey sample size](/docs/survey-sample-size-guide) for setting n in the first place.\n\n## Segment norms beat overall norms\n\nAn overall 4.1 can be 4.5 among enterprise customers and 3.4 among SMB, and the overall figure will move whenever your customer mix moves — even when nothing about either group's experience has changed. That is a mix shift masquerading as a finding, and it is one of the most common false alarms in tracking research.\n\nBuild norms per segment for the segments you actually make decisions about, and enforce your minimum n *within* each segment rather than overall. A segment with 30 responses does not get a percentile rank; it gets a range and a caveat. If your mix changes materially between waves, [weighting](/docs/survey-weighting-guide) lets you compare like with like.\n\n## When you have no history at all\n\nEveryone starts here. Three moves make the first year productive rather than wasted:\n\n1. **Declare wave one as the baseline explicitly.** Not as a judgement, as a reference point. Say so in the report: \"this is the baseline; no percentile interpretation is available until we have a distribution.\"\n2. **Use criterion-referencing in the interim** — but derive the criterion rather than inventing it. The most defensible source is qualitative: ask people what would make it a 5.\n3. **Ask why, from the first wave.** A norm bank tells you a score is unusual. It never tells you what caused it. If you only start collecting the reasoning once the number moves, you will have a distribution and no explanation.\n\nThat third point is where the design of your instrument matters more than the statistics. A tracking survey gives you a number and a shrug. A conversational interview gives you the number *and* the reason, from the same participant, in the same sitting.\n\n## How Koji makes a norm bank practical\n\nThe hardest part of norm-referencing is not the arithmetic — it is holding the instrument still for two years while stakeholders ask to reword things, and capturing the explanation alongside the number.\n\n- **[Structured questions](/docs/structured-questions-guide) freeze the instrument.** All six types — `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no` — are defined explicitly, and a `scale` question carries a fixed minimum, maximum and labels. Wave twelve asks exactly what wave one asked, and the response distribution is stored as discrete values rather than reconstructed from prose. That is a norm bank you can actually trust.\n- **AI follow-up questions capture the why at the moment the number is given.** When a respondent selects a 3, Koji's AI asks what would have made it higher — so every wave in your bank arrives with its own explanation attached. No industry benchmark table can tell you why your score moved; your own participants can, and they will if something asks them.\n- **Voice and text interviews** hold the mode constant across waves while still letting participants choose how to answer — and mode consistency is the single largest threat to comparability in a tracking programme.\n- **Real-time reports** recompute themes, quotes and quantitative summaries as interviews complete, so a wave closes with its distribution ready to store rather than waiting on manual analysis.\n\nThe result is that the tedious discipline norm-referencing demands — same instrument, stored distributions, reasoning captured alongside the number — becomes a property of how the study runs rather than a process somebody has to remember.\n\n## Frequently asked questions\n\n**How many past waves do I need before percentile ranks mean anything?**\nAround 8–10 as a working minimum, and 12 or more for comfortable interpretation. Below that, each wave carries too much percentile weight — with five waves in the bank, one wave moves the answer by 20 points. Report the range and the median instead until the bank is deep enough.\n\n**Should I use an industry benchmark or my own history?**\nYour own history, whenever you have enough of it, because it is the only comparison where the method is guaranteed to match. Use industry benchmarks for orientation when entering a new category or setting an initial target, and always check the mode, scale, wording and sample source behind the published figure before you compare against it.\n\n**What is the percentile rank formula for a norm bank?**\nPercentile rank = (B + 0.5E) / N × 100, where B is the number of past waves below the current score, E is the number exactly equal, and N is the total number of waves. The half-weighting of ties is what stops identical scores producing different percentiles depending on ordering.\n\n**Why did our score drop when nothing changed?**\nCheck the mix before you check the experience. If the proportion of responses from a lower-scoring segment increased, the overall score falls without any group's experience changing. This is why segment-level norms matter, and why weighting exists. Also check whether the drop exceeds your minimum detectable effect — roughly twice your margin of error for a wave-over-wave comparison — before treating it as real at all.\n\n**Should our headline metric be the mean or top-box?**\nDecide before you see the data, and store the full response distribution either way so you can compute both. Top-box is more sensitive to real change and noisier; the mean is more stable and compresses movement. Moving 10% of respondents from a 4 to a 5 shifts the mean by 0.10, the top-box by 10 points, and the top-two-box by nothing at all.\n\n**Can I add old waves to a norm bank if the wording changed?**\nNo. A wording, scale or mode change breaks comparability, and including pre-change waves widens your distribution with results that were never measuring the same thing. Start a new bank at the change, note the break in your documentation, and keep the old bank for historical reference only.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) — the six question types that keep an instrument fixed across waves\n- [Statistical Power and Minimum Detectable Effect](/docs/statistical-power-minimum-detectable-effect) — the arithmetic behind the minimum n for a usable norm bank\n- [System Usability Scale (SUS)](/docs/system-usability-scale-guide) — the best worked example of a properly normed instrument\n- [NPS Benchmarks by Industry](/docs/nps-benchmarks-by-industry-2026) — an industry table, and how to check it before using it\n- [Customer Experience Benchmarking](/docs/customer-experience-benchmarking) — measuring against external standards when they genuinely fit\n- [Brand Tracking Studies](/docs/brand-tracking-study-guide) — running the wave programme a norm bank is built from\n- [Survey Weighting: How to Correct a Skewed Sample](/docs/survey-weighting-guide) — comparing like with like when your mix shifts\n\n---\n\n*Want every wave to arrive with its distribution and its explanation? [Start free with 10 credits](https://www.koji.so) and run a tracking study where the AI asks why the number moved.*\n- [Every Input Was Accurate and the Difference Was Not](/docs/catastrophic-cancellation-metric-differences) — why the gap to your norm is less precise than the norm itself.\n- [Credibility weighting for small segment estimates](/docs/credibility-weighting-small-segment-estimates) — related reading\n- [How much data a segment needs before its own number is enough](/docs/full-credibility-standard-sample-size-per-segment) — related reading\n","category":"Research Methods","lastModified":"2026-09-18T12:22:19.838625+00:00","metaTitle":"Internal Benchmarks and Percentile Norms: Is Your Score Good? (2026)","metaDescription":"Build a norm bank from your own history instead of borrowing a mismatched industry benchmark. Percentile rank arithmetic, top-box vs mean, and the sample size below which it is all noise.","keywords":["internal benchmarks","percentile rank survey","norm-referencing research","is my score good","top-box scoring","norm bank tracking study","benchmark alternatives"],"aiSummary":"When no published industry benchmark matches your instrument, build a norm bank: a stored distribution of your own past waves against which new scores become percentile ranks. Published benchmarks mislead because of mode effects, scale and wording differences, sample source, question order, industry mix and self-selection in the benchmark itself - a benchmark is comparable only if the method matches, not the metric name. The Sauro-Lewis SUS curved grading scale, built from 241 studies where 68 equals a C at the 50th percentile, works because the instrument is fixed. Seven steps: freeze the instrument, log every wave with metadata, set a minimum n, store the full distribution not just the mean, compute percentile rank as (B + 0.5E)/N x 100, re-baseline on a schedule, publish the rules. Waves under n=100 do not belong in the bank and fewer than 8-10 waves should be reported as a range; minimum detectable effect is roughly 2.02 times margin of error for wave-over-wave comparison. Moving 10% of respondents from 4 to 5 shifts the mean 0.10, top-box 10 points and top-two-box zero.","aiPrerequisites":["A repeating survey or tracking programme with at least one prior wave","Basic familiarity with means, distributions and margin of error"],"aiLearningOutcomes":["Choose between criterion-, norm- and change-referenced interpretation","Judge whether a published industry benchmark is comparable to your measurement","Build and maintain a norm bank with the right metadata","Compute percentile ranks that handle ties correctly","Set minimum sample sizes below which percentile interpretation is noise","Choose between mean, top-box and top-two-box before seeing the data"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}