Back to docs
Analysis & Synthesis

Number Needed to Treat: How Many Users You Must Reach to Keep One (2026)

Every effect in your deck is a rate. None of them is a count of people. Number needed to treat converts a percentage lift into the only figure a roadmap can cost, and the evidence says the persuasive format is the misleading one.

Answer first: every effect in your product deck is expressed as a rate or a percentage lift, and almost none of them are expressed as a count of people. Number needed to treat closes that gap - it is one divided by the absolute risk reduction, and it answers the only question a roadmap actually needs answered: how many users do we have to reach to produce one more of the outcome we want? A "25 percent reduction in churn" and "you must fix onboarding for 34 accounts to save one" are the same finding. Only one of them can be costed.

Medicine solved this thirty-eight years ago and product management never picked it up.

Andreas Laupacis, David Sackett and Robin Roberts introduced the measure in the New England Journal of Medicine (318(26):1728-1733, 1988), and the idea is almost insultingly simple. If an intervention drops the probability of a bad outcome from 12 percent to 9 percent, the absolute risk reduction (ARR) is 3 percentage points. Its reciprocal, 1 / 0.03 = 33.3, is the number needed to treat: on average you must apply the intervention to about 34 people to prevent one bad outcome. Thirty-three of them get the intervention and no benefit, because they were never going to have the outcome or were going to have it anyway.

That last sentence is the reason the measure exists, and the reason nobody wants to put it on a slide.

The same result, four ways

Take a plausible product finding: a redesigned onboarding flow, tested properly, reduces 90-day cancellation among new accounts from 12 percent to 9 percent.

FormatThe numberWhat it invites you to concludeWhat it hides
Relative risk reduction"Cuts 90-day churn by 25 percent"This is a large effectThe baseline. A 25 percent cut of 0.4 percent is nothing
Absolute risk reduction"Cuts 90-day churn by 3 percentage points"This is a modest, real effectHow much work is needed to collect it
Number needed to treat"Fix onboarding for 34 accounts to save 1"This costs something per unit of outcomeNothing much - this is the honest form
Expected annual outcome"400 accounts per quarter, so about 48 saves a year"This is what it is worthNothing, provided the baseline is right

All four rows describe one experiment. The first one is the row that gets presented. It is also, measurably, the row that makes the effect look biggest.

The evidence that the persuasive format is the misleading one

This is not a hunch. It is one of the better-replicated findings in the literature on communicating evidence.

Naylor, Chen and Strauss ran the canonical test at Toronto teaching hospitals and published it in Annals of Internal Medicine (117(11):916-921, 1992). They took real endpoints from the Helsinki Heart Study, randomly assigned 100 faculty and housestaff to one of two questionnaires - one presenting results as absolute differences, one as relative differences - and asked each respondent to rate therapeutic effectiveness on an 11-point scale. Ratings from the 50 clinicians who saw absolute event data were lower than those from the 50 who saw relative risk reductions (P < 0.001), though the size of that gap was modest: no endpoint differed by more than 0.6 scale points.

The dramatic result was the fourth item. When the same data were presented as "77 persons treated for five years to prevent one myocardial infarction" - a number needed to treat, in other words - mean ratings fell by 2.3 scale points relative to the relative-risk framing and 1.8 points relative to the absolute-risk framing (both P < 0.001). Four times the effect of the relative-versus-absolute switch. The authors concluded that clinicians' views of therapies are shaped by the common use of relative risk reductions "in both trial reports and advertisements," by which endpoint gets emphasised, and "above all, by underuse of summary measures that relate treatment burden to therapeutic yields in a clinically relevant manner."

Substitute "roadmap documents and vendor case studies" for "trial reports and advertisements" and the sentence needs no other edit.

The Cochrane systematic review by Akl and colleagues (Cochrane Database of Systematic Reviews, 2011, Issue 3) pooled 35 studies reporting 83 comparisons across health professionals and consumers, and quantified the trade-off:

ComparisonUnderstandingPerceived size of effectPersuasiveness
Relative risk reduction vs absolute risk reductionNo real difference (SMD 0.02)RRR perceived larger (SMD 0.41)RRR more persuasive (SMD 0.66)
Relative risk reduction vs number needed to treatRRR better understood (SMD 0.73)RRR perceived larger (SMD 1.15)RRR more persuasive (SMD 0.65)
Absolute risk reduction vs number needed to treatARR better understood (SMD 0.42)ARR perceived larger (SMD 0.79)No real difference (SMD 0.05)

Their conclusion, verbatim: "Relative risk reduction (RRR), compared with absolute risk reduction (ARR) and number needed to treat (NNT), may be perceived to be larger and is more likely to be persuasive."

Read the table honestly and it says two things at once, and you need both. The format that persuades your stakeholders is the format that makes the effect look biggest. And NNT is genuinely harder to understand than either alternative - that is a real finding, not a slur, and it is why NNT belongs alongside the absolute number rather than instead of it. The practical rule that falls out: lead with the absolute risk reduction, put the count next to it, and keep the relative figure for the appendix where it can be checked against the baseline.

And almost nobody reports it

Nuovo, Melnikow and Chang reviewed every issue of Annals of Internal Medicine, the BMJ, JAMA, The Lancet, and the New England Journal of Medicine for four separate years, and identified 359 randomised controlled trials of a medication showing a significant treatment effect. NNT was reported in 8 of the 359. Absolute risk reduction was reported in 18 (JAMA 287:2813-2814, 2002).

Roughly two percent, in the five most rigorous journals in medicine, under the CONSORT reporting standard. If that is the base rate where reviewers are paid to check, the base rate in a product review deck is not a mystery.

Why the count is the only figure you can act on

A percentage lift cannot be compared to anything. A count can be compared to everything - to engineering weeks, to CSM hours, to a support budget, to the alternative on the roadmap.

Every intervention has a per-unit reach cost. A redesigned onboarding flow costs engineering time once and then reaches everyone for free, so a large NNT is fine. A human-led onboarding call costs a real hour per account, so an NNT of 34 means 34 hours per retained account, and that number can be put next to the account's annual value in a single line. The same experimental result therefore justifies one intervention and kills the other - and you cannot see that at all while the finding is expressed as "a 25 percent reduction."

The mirror measure is number needed to harm (NNH), and it is the discipline that separates a real recommendation from an enthusiastic one. If your mandatory onboarding checklist prevents one cancellation per 34 accounts but causes one additional signup abandonment per 100, you have an NNT of 34 against an NNH of 100 - roughly three saves per abandonment, which is probably worth it. If the abandonment rate is one in 25, you are destroying value while reporting a win. Almost no product experiment is analysed for its harm arm, because the harm usually lands on a different metric owned by a different team.

Three things that will make your NNT wrong

1. A baseline computed on the wrong event. NNT is 1 / (baseline risk x relative reduction), so the baseline does all the work. If your baseline churn probability came from a retention curve that treated acquisitions, downgrades and payment failures as censored rather than as competing risks, the baseline is inflated - and an inflated baseline makes NNT look smaller, which is to say makes your programme look cheaper than it is. In the worked medical example behind that article, a naive baseline of 43.0 percent against a correct 36.8 percent would understate NNT by about 15 percent across the board. Every save programme costed off that baseline is under-budgeted by the same margin.

2. An intervention that does not reach the people you counted. NNT assumes the treated were treated. If your "treatment" is an in-app prompt with a 40 percent open rate, the number you should report is the NNT among those exposed, and separately the number needed to offer - which is 2.5 times larger. Reporting the first while budgeting for the second is one of the more common ways a growth programme misses its forecast.

3. A time horizon left implicit. NNT is meaningless without one. "34 accounts to save one" over 90 days and over three years are wildly different claims. Naylor's respondents were told "treated for five years." Say your window out loud, and make sure it is a window your experiment actually observed rather than one your model extrapolated to.

Where research produces the number

Two of the three inputs to an NNT are quantitative and already in your warehouse: baseline risk and treated risk. The third input is not a number at all, and it is the one that decides whether the calculation is worth doing.

You need to know what the intervention actually did to the people it worked on. An NNT of 34 tells you 33 accounts received something with no effect. Some of those 33 were never at risk. Some were beyond saving. Some were at risk and the intervention simply did not address their reason. Those three groups have completely different implications: the first says target better, the second says accept a ceiling, the third says the intervention is aimed at the wrong problem. The arithmetic cannot tell them apart. Interviews can - and interviewing the non-responders is the study almost nobody runs, because the instinct is to interview the successes.

This is also the point where the two preceding articles in this cluster converge. Competing risks tells you which baseline you are entitled to use. Time-to-repair tells you that a save programme aimed at prevention may be competing against a much cheaper intervention aimed at recovery. NNT is the common currency in which those two options can finally be compared, because it converts both of them into the same unit: people you must reach per outcome you gain.

How Koji helps

Getting to the non-responder story used to be the blocker. It requires talking to the accounts where the intervention did nothing, which is a low-status, low-enthusiasm study that never wins a research prioritisation debate against a shiny discovery project.

Koji makes it a background process rather than a project:

  • AI-moderated interviews run across the entire treated cohort, not a convenience sample of the successes. If 400 accounts entered the programme, you can interview all 400 rather than the 12 whose CSM had time.
  • Structured questions turn the classification into typed data you can join back to the experiment. All six types are available - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and the non-responder triage is a yes_no on whether they noticed the intervention at all, a single_choice on which of the three non-response reasons applies, a scale on how close they came to leaving anyway, a ranking of what would have mattered more, a multiple_choice on what they tried instead, and an open_ended probe that the AI follows wherever the earlier answers point.
  • Automatic thematic analysis segments themes by outcome arm, so "why it worked" and "why it did nothing" are two theme lists rather than one blended one. Blending them is how a genuinely mis-aimed intervention survives a post-mortem.
  • Customizable AI consultants let you hold the interview protocol constant across quarters, which is what makes an NNT comparable over time rather than a one-off number.
  • Real-time reporting puts the exposure rate - the share who actually received the treatment - in front of you while the experiment is running, which is the input most often assumed rather than measured.

Compared with a legacy stack, the difference is which population you can afford to study. A traditional survey tool can send a questionnaire to everyone, but it cannot ask a follow-up, so a non-responder who says "I did not notice it" produces a dead end instead of a diagnosis. A traditional interview programme can probe properly, but at five to eight interviews a week it will never cover a treated cohort of 400 while the result is still relevant. AI moderation removes the trade-off: full-cohort coverage with adaptive follow-up, in days.

A working checklist

  1. For every effect you report, compute the absolute risk reduction and its reciprocal. Two lines of arithmetic.
  2. State the time horizon in the same sentence as the count. Always.
  3. Report the exposure rate separately, and give both the number needed to treat and the number needed to offer.
  4. Compute the number needed to harm on the metric most likely to move against you, and name who owns that metric.
  5. Confirm the baseline came from an event-specific estimate, not a blended attrition curve.
  6. Interview the non-responders, and split them into never-at-risk, beyond-saving, and wrong-problem.
  7. Put the relative figure in the appendix. It is the number most likely to be quoted back at you, and the one least likely to survive a check against the baseline.

Frequently asked questions

What is number needed to treat in a product context?

It is the number of users or accounts you must apply an intervention to in order to produce one additional instance of the outcome you want. It equals one divided by the absolute risk reduction. If a change reduces 90-day cancellation from 12 percent to 9 percent, the absolute risk reduction is 3 percentage points and the number needed to treat is about 34 accounts per save.

How is it different from a percentage lift?

A percentage lift is usually a relative risk reduction, which is the change divided by the baseline. It is the same result expressed without the baseline, so a 25 percent reduction can describe a 3-point change or a 0.1-point change. Number needed to treat forces the baseline back into the number, which is why it can be costed and a relative lift cannot.

Why do stakeholders find relative numbers more convincing?

Because they are. The Cochrane review by Akl and colleagues (2011) pooled 35 studies and 83 comparisons and found relative risk reduction was perceived as a larger effect and was more persuasive than either absolute risk reduction or number needed to treat. That is precisely why an evidence-led team should present the absolute figure first - the persuasive format and the accurate format are not the same format.

Is number needed to treat harder to understand?

Yes, and the evidence says so plainly: in the Cochrane pooled comparisons, both relative and absolute risk reductions were better understood than the number needed to treat. Use it as a companion to the absolute risk reduction rather than a replacement for it, and always attach the time horizon and the population.

What is number needed to harm?

The same calculation applied to an adverse outcome: how many users you must expose to the change to cause one additional instance of the bad thing. It is the discipline that stops a net-negative intervention from shipping with a positive headline, and it usually lives on a metric owned by a different team than the one running the experiment.

How many trials actually report this?

Very few. Nuovo, Melnikow and Chang reviewed 359 randomised controlled trials with significant treatment effects across the five most cited medical journals and found number needed to treat reported in 8 of them and absolute risk reduction in 18 (JAMA, 2002). If the reporting rate is about two percent under peer review and the CONSORT standard, expect it to be lower in a product review.

Related Resources

Related Articles

Competing Risks: Why Your Retention Curve Overstates the Churn You Care About (2026)

Your retention curve treats acquisitions, downgrades and payment failures as if those accounts were still at risk of cancelling. That inflates the number. Here is the correction, the size of the error, and the interview that produces the missing field.

Availability, Not Uptime: Why Time-to-Repair Is Half Your Retention Equation (2026)

Availability is a ratio with two terms, and product teams fund only one of them. Halving repair time and halving failure rate produce exactly the same result. Here is the arithmetic, the invisible parts of the customer repair clock, and how to measure them.

Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)

Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.

Statistical Significance in Survey Research: A Plain-English Guide (2026)

A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survey Sample Size: How Many Responses Do You Really Need? (2026 Guide)

A practical guide to survey sample size — formulas, calculators, real benchmarks by use case, and why AI-moderated interviews change the qual-vs-quant tradeoff entirely.

How to Write a User Research Report: Structure, Templates, and Best Practices

Learn how to structure a user research report that drives decisions — covering executive summary, key findings, data visualizations, themes, and recommendations. Includes how Koji generates reports automatically as interviews complete.