Back to docs
Research Methods

Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)

How to rate the severity of usability problems using Nielsen's 0-4 scale, why single-evaluator ratings are unreliable, how to separate severity from priority, and how to replace guessed frequency estimates with measured data.

Usability Issue Severity Ratings: How to Score, Prioritize, and Report UX Problems (2026)

Short answer: A severity rating scores how badly a usability problem hurts users, combining frequency, impact, persistence, and market impact. The standard is Jakob Nielsen's 0–4 scale, from "not a problem" to "usability catastrophe." The critical caveat is that severity ratings are judgments, and judgments vary enormously between evaluators — Nielsen is explicit that "severity ratings from a single evaluator are too unreliable to be trusted," and recommends averaging ratings from three evaluators. Severity is not the same thing as priority: severity measures user harm, priority weighs harm against reach and cost to fix.

Every usability study ends the same way. You have 47 observations, a stakeholder with capacity for four, and a meeting in which the loudest opinion wins. Severity ratings exist to replace that meeting with something defensible.

They are also the single most casually applied technique in UX. Most teams assign severity in a rush at the end of synthesis, by one person, using an undefined scale, and then present the numbers as if they were measurements. The evidence says that is closer to a coin flip than most practitioners believe — and the fix is neither complicated nor expensive.

The standard scale: Nielsen's 0–4

The reference scale comes from Jakob Nielsen's severity ratings work at Nielsen Norman Group:

RatingMeaningImplied action
0Not a usability problem at allDrop it; log the disagreement
1Cosmetic problem onlyFix only if spare time exists
2Minor usability problemLow priority
3Major usability problemImportant to fix; high priority
4Usability catastropheImperative to fix before release

The 0 rating is not filler. It is the escape hatch that lets a rater say "I do not agree this is a problem," and the volume of 0s in your data is a direct measure of how much your findings list is padded with observations that are actually preferences.

What severity actually combines

Severity is not a single dimension. Nielsen defines it as a combination of four factors:

FactorQuestionCommon failure
FrequencyIs this common or rare?Estimated from five participants and reported as fact
ImpactCan users overcome it easily or not?Confused with how annoying the evaluator found it
PersistenceOne-time hurdle, or does it bite repeatedly?Ignored entirely — the most under-weighted factor
Market impactDoes it affect the product's appeal and adoption?Either forgotten, or used to smuggle in business opinion

Persistence deserves special attention because it inverts intuitions. A confusing first-run flow that users solve once and never think about again is often rated 4 in the room and deserves a 2. A mildly awkward interaction on a screen someone visits forty times a day is often rated 2 and deserves a 4. Frequency of encounter and persistence of pain are different axes, and collapsing them is the most common severity error there is.

The uncomfortable evidence: the evaluator effect

Here is what the research says about how much you should trust a severity number.

Hertzum and Jacobsen's review of usability evaluation methods — the paper that named the evaluator effect — found that the average agreement between any two evaluators assessing the same system with the same method ranges from 5% to 65%, with no single method consistently outperforming the others. The effect appears both in which problems get detected and in how severely they get rated (Hertzum & Jacobsen, IJHCI 2003).

A more recent replication in unmoderated testing found the same pattern in practice. Four evaluators — ranging from about 100 sessions of experience to more than 10,000 — independently analysed the same unmoderated study. They logged 119 issue instances that consolidated to 38 unique problems. Pairwise agreement averaged 41%, ranging from 25% to 48%. Most strikingly, 47% of the unique problems were found by only one of the four evaluators, and only 18% were found by all four (MeasuringU).

Read that again: nearly half the findings in a typical study exist because one particular person watched the sessions. Change the analyst and you change the report.

This is not an argument against usability testing. It is an argument against single-rater severity, and Nielsen's guidance follows directly from it: the mean of ratings from three evaluators is satisfactory for practical purposes, and rating quality improves rapidly as raters are added.

Choose a scale you can actually apply

The 0–4 scale is the standard, but it is not the only defensible option, and there is a real case for fewer levels. MeasuringU reduced their own scale from seven points to three — Minor, Moderate, Critical, plus a separate non-problem category for insights and suggestions — after repeatedly finding the distinction between middle categories murky in practice.

ScaleWhen to use
Nielsen 0–4Formal reports, heuristic evaluation, regulated contexts, when you need a defensible standard
3-level (Minor / Moderate / Critical)Fast-cycle product teams; raters apply it more consistently
Binary blocker / non-blockerPre-release triage only; loses too much for research reporting

Whatever you pick, write the anchors down and give each level an observable definition, not an adjectival one. "Causes task failure" is observable. "Serious" is not. A rubric with observable anchors is the cheapest available improvement to inter-rater agreement — the same principle that governs inter-rater reliability in qualitative coding.

A workable set of anchors:

  • Critical (4): Participant could not complete the task, or completed it incorrectly without realising. Data loss, security exposure, or an accessibility barrier that excludes a user group entirely.
  • Major (3): Participant completed the task only after a workaround, backtracking, or help. Substantial time cost or visible frustration.
  • Minor (2): Participant hesitated, took a wrong turn, and self-corrected within seconds. Task succeeded.
  • Cosmetic (1): Noticed and commented on, but no effect on task performance.
  • Not a problem (0): The rater disagrees that this is a usability issue.

Severity is not priority

This is the distinction that determines whether your severity ratings survive the roadmap meeting.

Severity is a property of the problem's effect on users. It does not know or care what it costs to fix.

Priority is a business decision that combines severity with reach, effort, strategic value, and risk. A cosmetic issue on the checkout page that 100% of paying customers see can rationally outrank a catastrophe in an admin screen used by nine people once a quarter.

Keep them in separate columns. The moment you let fix-cost leak into the severity rating, you lose the ability to say "we knowingly shipped a major usability problem because the fix was expensive" — which is exactly the sentence that protects a research team's credibility when the issue resurfaces in support tickets six months later.

The clean handoff is: research owns severity, and the product team feeds severity into whatever prioritisation framework it already runs — RICE, ICE, or a value vs. effort matrix. Severity is an input to prioritisation, not a competitor to it.

A rating protocol that takes 45 minutes

  1. Compile the raw findings list with a one-line observable description of each problem, the participant IDs who hit it, and a timestamp or clip.
  2. Rate independently, in silence. Three raters minimum. Include at least one person who did not run the sessions — the evaluator effect is strongest among people who share a mental model.
  3. Never rate in a group first. Group rating produces consensus, not agreement. The first confident voice anchors everyone else, and you lose the disagreement signal that is the most useful output of the exercise.
  4. Compute the mean and the spread. Report both. A problem rated 4, 4, 4 is a different object from one rated 4, 3, 1, even though the second averages to a respectable 2.7.
  5. Discuss only the split items. Anything with a range of two or more points gets a five-minute conversation. Usually one rater knows something the others do not — a support-ticket volume, an accessibility implication, a known workaround. That knowledge is the actual finding.
  6. Record the rationale for anything rated 4. Catastrophes get challenged. Write down the frequency and impact evidence at rating time, not when someone questions it in a roadmap review.

Fix the weakest input: stop guessing frequency

Look back at the four factors. Two of them — impact and persistence — genuinely require expert judgment. But frequency is not a judgment. It is a measurement, and almost every team estimates it from five participants because measuring it properly used to be prohibitively expensive.

That is the real leverage point, and it is where a modern research stack changes the arithmetic.

Running a study through Koji lets you replace the guessed inputs with measured ones:

  • Frequency becomes an actual rate. Because AI-moderated sessions run asynchronously and in parallel, running 40 or 60 participants costs roughly what scheduling 8 used to. "3 of 5 participants" becomes "38% of 60 participants," and a severity rating built on that number survives scrutiny in a way the first one never does.
  • Impact gets measured from the participant, not inferred by the observer. Use Koji's structured questions directly in the flow: a scale question after each task captures perceived effort — the same logic behind the Single Ease Question — and a yes_no question captures self-reported success, which you can compare against observed success to catch the users who failed without knowing it. That gap is your highest-severity population.
  • Persistence gets asked instead of assumed. An open_ended question — "If you hit this again tomorrow, what would you do?" — with AI follow-up probing distinguishes a one-time hurdle from a recurring one in the participant's own words.
  • Relative harm comes from a ranking question. Have participants rank the problems they encountered. Evaluators are systematically bad at guessing which annoyance users actually care about, and a ranked aggregate is a direct corrective.
  • single_choice and multiple_choice let you segment severity by user type, which frequently reveals that a "minor" issue is a catastrophe for one segment.

Koji's automatic thematic analysis clusters the same problem across dozens of sessions so the frequency count is produced for you rather than tallied by hand, and every session carries a quality score on a 1–5 scale so you can tell which sessions actually carried signal. The point is not that AI replaces the rater. It is that the rater ends up judging two factors instead of four, with real numbers underneath the other two.

Reporting severity so it drives action

  • Lead with the 4s and 3s. Nobody reads a table of 47 rows. Put the catastrophes and majors above the fold with evidence clips.
  • Show the rater spread, not just the mean. Disagreement is information about how confident the team should be.
  • Attach evidence to every rating above 2. A timestamp, a quote, a clip. Severity claims without evidence get relitigated forever.
  • Report frequency as a rate with a denominator. "12 of 60" beats "several participants" every time.
  • Never present severity as a fix order. Hand it to the prioritisation framework and say so explicitly.

For the full report structure this fits into, see the UX research report guide.

Common mistakes

  1. One person rating everything. The evidence is unambiguous on this.
  2. Group rating before independent rating. You get consensus, and lose the disagreement signal.
  3. Letting fix cost into the severity score. Now you cannot separate "not bad" from "not worth fixing."
  4. Ignoring persistence. The most under-weighted of the four factors.
  5. Reporting the mean without the range. 4/3/1 and 3/3/2 are not the same finding.
  6. Guessing frequency from five participants and stating it as fact. Measure it.
  7. No 0 option. Without it, every observation becomes a problem, and your list inflates.

Frequently asked questions

What is the standard usability severity rating scale? Jakob Nielsen's 0–4 scale is the standard: 0 = not a usability problem, 1 = cosmetic, 2 = minor, 3 = major, 4 = usability catastrophe. Severity combines four factors — frequency, impact, persistence, and market impact. Some teams use a simpler three-level Minor / Moderate / Critical scale, which raters tend to apply more consistently because the middle distinctions in longer scales are hard to hold steady.

How many people should rate severity? At least three. Nielsen states that severity ratings from a single evaluator are too unreliable to be trusted, and that the mean of three evaluators' ratings is satisfactory for practical purposes, with quality improving rapidly as raters are added. Rate independently first, then discuss only the items where ratings diverge by two or more points.

What is the evaluator effect? It is the finding that different evaluators analysing the same sessions identify different problems and rate them differently. Hertzum and Jacobsen found average any-two agreement between evaluators ranging from 5% to 65% across methods. A replication in unmoderated testing found 41% average pairwise agreement, with 47% of unique problems detected by only one of four evaluators. It is the main reason single-rater severity should not be trusted.

What is the difference between severity and priority? Severity describes how much a problem harms users, independent of what fixing it costs. Priority is a business decision that combines severity with reach, effort, strategic value, and risk. Keep them in separate columns so you can knowingly defer a major issue without pretending it is minor — and feed severity into a prioritisation framework like RICE or ICE rather than treating it as a fix order.

How do I estimate frequency with only five participants? Honestly, you cannot with much confidence, and the correct response is to report it as a raw count with the denominator rather than a percentage. The better answer is to stop being limited to five. Asynchronous AI-moderated testing makes 40 to 60 participants practical at a cost close to what eight scheduled sessions used to require, which converts frequency from an estimate into a measurement.

Should severity ratings come from evaluators or from users? Both, for different factors. Impact and persistence benefit from expert judgment, because participants often cannot tell how much time they lost or whether a workaround will scale. Frequency should be measured, and relative harm is best captured by asking participants directly — a ranking question across the problems they encountered corrects for the fact that evaluators are poor at guessing which annoyances users genuinely care about.

Related Resources

Related Articles

AI Failure Mode Analysis: An FMEA Framework for AI Products (2026)

How to run Failure Mode and Effects Analysis (FMEA) on an AI product: the failure mode taxonomy, how to score severity, occurrence and detection when failures are probabilistic, and how user research supplies the numbers.

AI Guardrail Testing: How to Measure False Refusals and Over-Blocking with Real Users (2026)

Your safety layer has a false positive rate, and it is costing you users you never hear from. How to measure false refusal rate, run an over-blocking study, and tune guardrails against real user harm instead of vibes.

Heuristic Evaluation: The Complete UX Review Guide

Learn how to conduct heuristic evaluations using Nielsen's 10 usability heuristics. Discover when to use expert review vs. user testing, how many evaluators you need, and how AI-assisted research accelerates the process.

Inter-Rater Reliability in Qualitative Research: A Practical Guide to Coding Agreement

Learn how to measure inter-rater (intercoder) reliability in qualitative research using Cohen's kappa and Krippendorff's alpha, what thresholds count as reliable, and how AI-native tools make consistent coding the default.

RICE Prioritization Framework: How to Score and Rank Product Ideas

Master the RICE scoring framework (Reach, Impact, Confidence, Effort) for product prioritization. Includes the formula, worked examples, free template, and how customer research transforms Confidence scores.

Single Ease Question (SEQ): The 7-Point UX Metric for Task-Level Usability (2026)

The complete 2026 guide to the Single Ease Question (SEQ): the verbatim 7-point scale wording, Sauro–MeasuringU benchmarks (5.3–5.5 average), correlation with task completion, when to use SEQ vs SUS, and how to bundle SEQ into AI-moderated interviews on Koji to get task-level usability scores in days.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Usability Metrics: Task Success Rate, Time on Task, and Error Rate Explained

The complete guide to the core usability metrics — task success rate, time on task, and error rate — including industry benchmarks, formulas, sample sizes, and how to capture them automatically with AI-moderated research.

How to Conduct Usability Testing: The Complete Guide

A comprehensive guide to usability testing for UX researchers and product managers. Covers types of testing, participant numbers, step-by-step facilitation, and the most common mistakes to avoid.

How to Create Effective UX Research Reports (+ Free Template)

A complete guide to writing UX research reports that drive decisions — with a reusable template, best practices, and how AI tools like Koji auto-generate research reports in minutes.