{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-02T23:48:31.062Z"},"content":[{"type":"documentation","id":"dc164ea4-fdb0-4212-834b-3a0e841fdd29","slug":"usability-issue-severity-ratings","title":"Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)","url":"https://www.koji.so/docs/usability-issue-severity-ratings","summary":"Severity rates how badly a usability problem harms users, combining frequency, impact, persistence, and market impact on Nielsen's 0-4 scale. Single-evaluator ratings are unreliable: the evaluator effect produces 5-65% any-two agreement (Hertzum & Jacobsen), and one replication found 41% pairwise agreement with 47% of unique problems found by only one of four evaluators. Use three independent raters, report mean and spread, keep severity separate from priority, and measure frequency with larger asynchronous samples instead of estimating it from five participants.","content":"# Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)\n\n**Short answer:** A severity rating scores how badly a usability problem hurts users, combining **frequency, impact, persistence, and market impact**. The standard is Jakob Nielsen's 0–4 scale, from \"not a problem\" to \"usability catastrophe.\" The critical caveat is that severity ratings are judgments, and judgments vary enormously between evaluators — Nielsen is explicit that **\"severity ratings from a single evaluator are too unreliable to be trusted,\"** and recommends averaging ratings from three evaluators. Severity is not the same thing as priority: severity measures user harm, priority weighs harm against reach and cost to fix.\n\nEvery usability study ends the same way. You have 47 observations, a stakeholder with capacity for four, and a meeting in which the loudest opinion wins. Severity ratings exist to replace that meeting with something defensible.\n\nThey are also the single most casually applied technique in UX. Most teams assign severity in a rush at the end of synthesis, by one person, using an undefined scale, and then present the numbers as if they were measurements. The evidence says that is closer to a coin flip than most practitioners believe — and the fix is neither complicated nor expensive.\n\n## The standard scale: Nielsen's 0–4\n\nThe reference scale comes from Jakob Nielsen's severity ratings work at [Nielsen Norman Group](https://www.nngroup.com/articles/how-to-rate-the-severity-of-usability-problems/):\n\n| Rating | Meaning | Implied action |\n|---|---|---|\n| **0** | Not a usability problem at all | Drop it; log the disagreement |\n| **1** | Cosmetic problem only | Fix only if spare time exists |\n| **2** | Minor usability problem | Low priority |\n| **3** | Major usability problem | Important to fix; high priority |\n| **4** | Usability catastrophe | Imperative to fix before release |\n\nThe 0 rating is not filler. It is the escape hatch that lets a rater say \"I do not agree this is a problem,\" and the volume of 0s in your data is a direct measure of how much your findings list is padded with observations that are actually preferences.\n\n## What severity actually combines\n\nSeverity is not a single dimension. Nielsen defines it as a combination of four factors:\n\n| Factor | Question | Common failure |\n|---|---|---|\n| **Frequency** | Is this common or rare? | Estimated from five participants and reported as fact |\n| **Impact** | Can users overcome it easily or not? | Confused with how annoying the evaluator found it |\n| **Persistence** | One-time hurdle, or does it bite repeatedly? | Ignored entirely — the most under-weighted factor |\n| **Market impact** | Does it affect the product's appeal and adoption? | Either forgotten, or used to smuggle in business opinion |\n\nPersistence deserves special attention because it inverts intuitions. A confusing first-run flow that users solve once and never think about again is often rated 4 in the room and deserves a 2. A mildly awkward interaction on a screen someone visits forty times a day is often rated 2 and deserves a 4. Frequency of *encounter* and persistence of *pain* are different axes, and collapsing them is the most common severity error there is.\n\n## The uncomfortable evidence: the evaluator effect\n\nHere is what the research says about how much you should trust a severity number.\n\nHertzum and Jacobsen's review of usability evaluation methods — the paper that named the **evaluator effect** — found that the average agreement between any two evaluators assessing the same system with the same method ranges from **5% to 65%**, with no single method consistently outperforming the others. The effect appears both in *which* problems get detected and in *how severely* they get rated ([Hertzum & Jacobsen, IJHCI 2003](https://mortenhertzum.dk/publ/IJHCI2003.pdf)).\n\nA more recent replication in unmoderated testing found the same pattern in practice. Four evaluators — ranging from about 100 sessions of experience to more than 10,000 — independently analysed the same unmoderated study. They logged 119 issue instances that consolidated to **38 unique problems**. Pairwise agreement averaged **41%**, ranging from 25% to 48%. Most strikingly, **47% of the unique problems were found by only one of the four evaluators**, and only **18% were found by all four** ([MeasuringU](https://measuringu.com/examining-the-evaluator-effect-in-unmoderated-usability-testing/)).\n\nRead that again: nearly half the findings in a typical study exist because one particular person watched the sessions. Change the analyst and you change the report.\n\nThis is not an argument against usability testing. It is an argument against single-rater severity, and Nielsen's guidance follows directly from it: the mean of ratings from **three evaluators** is satisfactory for practical purposes, and rating quality improves rapidly as raters are added.\n\n## Choose a scale you can actually apply\n\nThe 0–4 scale is the standard, but it is not the only defensible option, and there is a real case for fewer levels. MeasuringU reduced their own scale from seven points to three — **Minor, Moderate, Critical**, plus a separate non-problem category for insights and suggestions — after repeatedly finding the distinction between middle categories murky in practice.\n\n| Scale | When to use |\n|---|---|\n| **Nielsen 0–4** | Formal reports, heuristic evaluation, regulated contexts, when you need a defensible standard |\n| **3-level (Minor / Moderate / Critical)** | Fast-cycle product teams; raters apply it more consistently |\n| **Binary blocker / non-blocker** | Pre-release triage only; loses too much for research reporting |\n\nWhatever you pick, **write the anchors down and give each level an observable definition**, not an adjectival one. \"Causes task failure\" is observable. \"Serious\" is not. A rubric with observable anchors is the cheapest available improvement to inter-rater agreement — the same principle that governs [inter-rater reliability in qualitative coding](/docs/inter-rater-reliability-qualitative-research).\n\nA workable set of anchors:\n\n- **Critical (4):** Participant could not complete the task, or completed it incorrectly without realising. Data loss, security exposure, or an accessibility barrier that excludes a user group entirely.\n- **Major (3):** Participant completed the task only after a workaround, backtracking, or help. Substantial time cost or visible frustration.\n- **Minor (2):** Participant hesitated, took a wrong turn, and self-corrected within seconds. Task succeeded.\n- **Cosmetic (1):** Noticed and commented on, but no effect on task performance.\n- **Not a problem (0):** The rater disagrees that this is a usability issue.\n\n## Severity is not priority\n\nThis is the distinction that determines whether your severity ratings survive the roadmap meeting.\n\n**Severity** is a property of the problem's effect on users. It does not know or care what it costs to fix.\n\n**Priority** is a business decision that combines severity with reach, effort, strategic value, and risk. A cosmetic issue on the checkout page that 100% of paying customers see can rationally outrank a catastrophe in an admin screen used by nine people once a quarter.\n\nKeep them in separate columns. The moment you let fix-cost leak into the severity rating, you lose the ability to say \"we knowingly shipped a major usability problem because the fix was expensive\" — which is exactly the sentence that protects a research team's credibility when the issue resurfaces in support tickets six months later.\n\nThe clean handoff is: research owns severity, and the product team feeds severity into whatever prioritisation framework it already runs — [RICE](/docs/rice-prioritization-framework), [ICE](/docs/ice-prioritization-framework), or a [value vs. effort matrix](/docs/value-vs-effort-prioritization-matrix). Severity is an input to prioritisation, not a competitor to it.\n\n## A rating protocol that takes 45 minutes\n\n1. **Compile the raw findings list** with a one-line observable description of each problem, the participant IDs who hit it, and a timestamp or clip.\n2. **Rate independently, in silence.** Three raters minimum. Include at least one person who did not run the sessions — the evaluator effect is strongest among people who share a mental model.\n3. **Never rate in a group first.** Group rating produces consensus, not agreement. The first confident voice anchors everyone else, and you lose the disagreement signal that is the most useful output of the exercise.\n4. **Compute the mean and the spread.** Report both. A problem rated 4, 4, 4 is a different object from one rated 4, 3, 1, even though the second averages to a respectable 2.7.\n5. **Discuss only the split items.** Anything with a range of two or more points gets a five-minute conversation. Usually one rater knows something the others do not — a support-ticket volume, an accessibility implication, a known workaround. That knowledge is the actual finding.\n6. **Record the rationale for anything rated 4.** Catastrophes get challenged. Write down the frequency and impact evidence at rating time, not when someone questions it in a roadmap review.\n\n## Fix the weakest input: stop guessing frequency\n\nLook back at the four factors. Two of them — impact and persistence — genuinely require expert judgment. But **frequency is not a judgment. It is a measurement**, and almost every team estimates it from five participants because measuring it properly used to be prohibitively expensive.\n\nThat is the real leverage point, and it is where a modern research stack changes the arithmetic.\n\nRunning a study through Koji lets you replace the guessed inputs with measured ones:\n\n- **Frequency becomes an actual rate.** Because [AI-moderated sessions](/docs/ai-usability-testing-guide) run asynchronously and in parallel, running 40 or 60 participants costs roughly what scheduling 8 used to. \"3 of 5 participants\" becomes \"38% of 60 participants,\" and a severity rating built on that number survives scrutiny in a way the first one never does.\n- **Impact gets measured from the participant, not inferred by the observer.** Use Koji's [structured questions](/docs/structured-questions-guide) directly in the flow: a `scale` question after each task captures perceived effort — the same logic behind the [Single Ease Question](/docs/single-ease-question-seq-guide) — and a `yes_no` question captures self-reported success, which you can compare against observed success to catch the users who failed without knowing it. That gap is your highest-severity population.\n- **Persistence gets asked instead of assumed.** An `open_ended` question — \"If you hit this again tomorrow, what would you do?\" — with AI follow-up probing distinguishes a one-time hurdle from a recurring one in the participant's own words.\n- **Relative harm comes from a `ranking` question.** Have participants rank the problems they encountered. Evaluators are systematically bad at guessing which annoyance users actually care about, and a ranked aggregate is a direct corrective.\n- **`single_choice` and `multiple_choice`** let you segment severity by user type, which frequently reveals that a \"minor\" issue is a catastrophe for one segment.\n\nKoji's automatic thematic analysis clusters the same problem across dozens of sessions so the frequency count is produced for you rather than tallied by hand, and every session carries a quality score on a 1–5 scale so you can tell which sessions actually carried signal. The point is not that AI replaces the rater. It is that the rater ends up judging two factors instead of four, with real numbers underneath the other two.\n\n## Reporting severity so it drives action\n\n- **Lead with the 4s and 3s.** Nobody reads a table of 47 rows. Put the catastrophes and majors above the fold with evidence clips.\n- **Show the rater spread**, not just the mean. Disagreement is information about how confident the team should be.\n- **Attach evidence to every rating above 2.** A timestamp, a quote, a clip. Severity claims without evidence get relitigated forever.\n- **Report frequency as a rate with a denominator.** \"12 of 60\" beats \"several participants\" every time.\n- **Never present severity as a fix order.** Hand it to the prioritisation framework and say so explicitly.\n\nFor the full report structure this fits into, see the [UX research report guide](/docs/ux-research-report-template).\n\n## Common mistakes\n\n1. **One person rating everything.** The evidence is unambiguous on this.\n2. **Group rating before independent rating.** You get consensus, and lose the disagreement signal.\n3. **Letting fix cost into the severity score.** Now you cannot separate \"not bad\" from \"not worth fixing.\"\n4. **Ignoring persistence.** The most under-weighted of the four factors.\n5. **Reporting the mean without the range.** 4/3/1 and 3/3/2 are not the same finding.\n6. **Guessing frequency from five participants and stating it as fact.** Measure it.\n7. **No 0 option.** Without it, every observation becomes a problem, and your list inflates.\n\n## Frequently asked questions\n\n**What is the standard usability severity rating scale?** Jakob Nielsen's 0–4 scale is the standard: 0 = not a usability problem, 1 = cosmetic, 2 = minor, 3 = major, 4 = usability catastrophe. Severity combines four factors — frequency, impact, persistence, and market impact. Some teams use a simpler three-level Minor / Moderate / Critical scale, which raters tend to apply more consistently because the middle distinctions in longer scales are hard to hold steady.\n\n**How many people should rate severity?** At least three. Nielsen states that severity ratings from a single evaluator are too unreliable to be trusted, and that the mean of three evaluators' ratings is satisfactory for practical purposes, with quality improving rapidly as raters are added. Rate independently first, then discuss only the items where ratings diverge by two or more points.\n\n**What is the evaluator effect?** It is the finding that different evaluators analysing the same sessions identify different problems and rate them differently. Hertzum and Jacobsen found average any-two agreement between evaluators ranging from 5% to 65% across methods. A replication in unmoderated testing found 41% average pairwise agreement, with 47% of unique problems detected by only one of four evaluators. It is the main reason single-rater severity should not be trusted.\n\n**What is the difference between severity and priority?** Severity describes how much a problem harms users, independent of what fixing it costs. Priority is a business decision that combines severity with reach, effort, strategic value, and risk. Keep them in separate columns so you can knowingly defer a major issue without pretending it is minor — and feed severity into a prioritisation framework like RICE or ICE rather than treating it as a fix order.\n\n**How do I estimate frequency with only five participants?** Honestly, you cannot with much confidence, and the correct response is to report it as a raw count with the denominator rather than a percentage. The better answer is to stop being limited to five. Asynchronous AI-moderated testing makes 40 to 60 participants practical at a cost close to what eight scheduled sessions used to require, which converts frequency from an estimate into a measurement.\n\n**Should severity ratings come from evaluators or from users?** Both, for different factors. Impact and persistence benefit from expert judgment, because participants often cannot tell how much time they lost or whether a workaround will scale. Frequency should be measured, and relative harm is best captured by asking participants directly — a ranking question across the problems they encountered corrects for the fact that evaluators are poor at guessing which annoyances users genuinely care about.\n\n## Related Resources\n\n- [Structured Questions Guide](/docs/structured-questions-guide) — the six question types for measuring impact and frequency\n- [How to Conduct Usability Testing](/docs/usability-testing-guide) — the complete method this fits into\n- [Usability Metrics: Task Success, Time on Task, and Error Rate](/docs/usability-metrics-guide) — the quantitative companions to severity\n- [Heuristic Evaluation Guide](/docs/heuristic-evaluation-guide) — where severity ratings are most commonly applied\n- [Single Ease Question (SEQ)](/docs/single-ease-question-seq-guide) — measuring perceived task difficulty\n- [Inter-Rater Reliability in Qualitative Research](/docs/inter-rater-reliability-qualitative-research) — improving agreement between raters\n- [RICE Prioritization Framework](/docs/rice-prioritization-framework) — where severity goes once research hands it off\n- [UX Research Report Template](/docs/ux-research-report-template) — reporting severity so it drives action\n","category":"Research Methods","lastModified":"2026-08-02T03:20:44.868648+00:00","metaTitle":"Usability Severity Ratings: Nielsen's 0-4 Scale & How to Apply It (2026)","metaDescription":"How to rate usability problem severity with Nielsen's 0-4 scale, why single-evaluator ratings are unreliable (41% agreement), how to separate severity from priority, and how to measure frequency instead of guessing it.","keywords":["usability severity ratings","severity rating scale usability","how to prioritize usability issues","nielsen severity scale","ux issue severity","usability problem severity","evaluator effect","severity vs priority"],"aiSummary":"Severity rates how badly a usability problem harms users, combining frequency, impact, persistence, and market impact on Nielsen's 0-4 scale. Single-evaluator ratings are unreliable: the evaluator effect produces 5-65% any-two agreement (Hertzum & Jacobsen), and one replication found 41% pairwise agreement with 47% of unique problems found by only one of four evaluators. Use three independent raters, report mean and spread, keep severity separate from priority, and measure frequency with larger asynchronous samples instead of estimating it from five participants.","aiPrerequisites":["Familiarity with usability testing or heuristic evaluation","A completed study with a raw findings list"],"aiLearningOutcomes":["Apply Nielsen's 0-4 severity scale with observable anchors","Explain the evaluator effect and why single-rater severity is unreliable","Run an independent-then-calibrate severity rating session with three raters","Separate severity from priority and hand off cleanly to a prioritisation framework","Replace guessed frequency estimates with measured rates from larger samples"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}