The Churn Hazard Curve: Why One Churn Rate Hides Three Different Problems (2026)
Your monthly churn rate averages three unrelated problems into one number. Learn to plot the churn hazard by tenure, read the three regimes, and avoid the sorting trap that makes a flattening curve look like product-market fit.
The conditional probability that an account churns in its ninth month is a different number, with a different cause and a different fix, than the conditional probability that it churns in its first. Your monthly churn rate averages them into one figure, and every intervention you fund off that figure is aimed at an average that describes no account in your book. The fix is to plot the hazard - churn as a share of accounts that survived to reach that tenure - instead of the retention curve, and then to run a separate study for each regime the hazard reveals.
Aviation learned this in the 1960s. When United Airlines plotted conditional failure probability against operating age for its components, the familiar wear-out story turned out to describe 6 percent of the fleet. This guide ports the analysis, the arithmetic, and the trap that comes with it.
Retention curve and hazard curve are not the same chart
Most teams have a retention curve. Very few have a hazard curve, and the two answer different questions.
A retention curve is a survival function: of the accounts that started, what share is still here at month t? It only ever goes down. A hazard curve is conditional: of the accounts that reached month t, what share left during month t? It can go down, stay flat, or go up.
Reliability engineering defines the second one precisely. It is "the probability that an item entering a given age interval will fail during that interval" - a measure also known as "the hazard rate or the local failure rate" (Nowlan and Heap, Reliability-Centered Maintenance, 1978, p. 25).
The arithmetic is a division you are probably not doing. In the United Airlines analysis of the Pratt and Whitney JT8D-7 engine, an engine had a probability of .692 of reaching 1,000 hours. The probability density of failure in the following 200-hour interval was .053. So the conditional probability of failure for an engine that actually reached 1,000 hours was .053 / .692 = .077 - noticeably higher than the unconditional .053.
Translate that to accounts:
| Month | Accounts entering the month | Churned during the month | Unconditional (share of original cohort) | Hazard (share of survivors) |
|---|---|---|---|---|
| 1 | 1,000 | 120 | 12.0% | 12.0% |
| 2 | 880 | 70 | 7.0% | 8.0% |
| 6 | 640 | 19 | 1.9% | 3.0% |
| 12 | 545 | 16 | 1.6% | 2.9% |
| 24 | 470 | 24 | 2.4% | 5.1% |
Read the fourth column and the story is "churn keeps falling, our curve is flattening, we have product-market fit." Read the fifth and there is a spike at month 1, a long quiet stretch, and something turning back up at month 24 that the flattening curve cannot show you, because a survival curve that declines slowly looks reassuring whether the underlying risk is falling or climbing.
Our cohort analysis guide covers the survival side properly - the three curve shapes and what a plateau implies about your market. This article is the derivative view, and it is the one that tells you what to fund.
What happened when an airline plotted the hazard instead
Nowlan and Heap's study is the most-cited empirical result in maintenance engineering, and the finding is not what anyone expected. Every item United analyzed fell into one of six age-reliability patterns:
| Pattern | Shape | Share of items |
|---|---|---|
| A | Bathtub: infant mortality, then flat, then pronounced wear-out | 4% |
| B | Flat or slowly rising, then pronounced wear-out | 2% |
| C | Gradually increasing, no identifiable wear-out age | 5% |
| D | Low when new, quick rise to a constant level | 7% |
| E | Constant at all ages (exponential) | 14% |
| F | Infant mortality, then constant or very slowly increasing | 68% |
The conclusion, verbatim: "Some 89 percent of the items analyzed had no wearout zone; therefore their performance could not be improved by the imposition of an age limit." Only 11 percent (patterns A, B and C) might benefit from a limit on operating age at all, and "Only 6% of the items studied showed pronounced wearout characteristics."
The most quoted line is about the shape everyone assumes is universal: "Although it is often assumed that the bathtub curve is representative of most items, note that just 4% of the items fell into this pattern."
Pattern F - 68 percent, the single largest group - is infant mortality followed by a flat tail, and Nowlan and Heap note it is "particularly applicable to electronic equipment." Complex items with many interacting parts do not wear out. They either fail early because something about the installation was wrong, or they run indefinitely. A B2B software account is a complex item.
The three regimes, translated
| Regime | What the hazard does | What is actually happening | The research question |
|---|---|---|---|
| Infant mortality (months 0-3) | High, falling fast | Wrong-fit accounts, failed implementation, the buyer's problem was misdiagnosed at sale | Why did this account never reach the value it was sold? |
| Useful life (months 4-18) | Low, roughly flat | Exogenous shocks: budget cuts, the champion leaves, reorgs, acquisition | What happened outside the product? |
| Wear-out (month 18+) | Rising again | Accumulated workarounds, outgrown the data model, a competitor closed the gap | What did we stop being able to do for you? |
These are three different studies with three different recruiting screens. A single "why did you churn" survey blended across all tenures produces an average of three unrelated causal stories, which is why the answers so often read as bland.
The regimes also rank your options. Infant-mortality churn is largely a sales-qualification and onboarding problem, and it is the cheapest to fix because the accounts are numerous and the cause is usually recent and recallable. Useful-life churn is mostly not addressable by the product at all - it is the constant-hazard background rate, and spending against it has poor returns. Wear-out churn is the expensive kind: the accounts are large, tenured, and the cause accumulated over a year before anyone noticed.
The trap: a falling hazard may be sorting, not loyalty
Here is the part that will change how you read your own chart, and it is the reason a falling hazard is weaker evidence than it looks.
Suppose every single customer has a fixed, unchanging monthly churn probability, but customers differ from each other - some are 1 percent per month, some are 20 percent. Nobody becomes more loyal over time. What does the aggregate hazard curve do?
It falls. Steeply.
The high-risk customers leave first, so the surviving population is progressively enriched with low-risk customers. The aggregate rate drops even though no individual's rate moved. Fader and Hardie make exactly this point about subscription businesses: increasing cohort-level retention "is purely due to cross-sectional heterogeneity, with individual customers having a constant propensity to churn. Cohort-level retention rates increase because those customers with high churn propensities drop out early on, leaving an ever-increasing proportion of customers who have low propensities to churn" (Fader, Hardie, Liu, Davin and Steenburgh, Journal of Interactive Marketing, 2018).
They note this "flies in the face of conventional wisdom, which assumes that a customer's propensity to churn decreases the longer their tenure with the firm." It is a known statistical result that "unobserved heterogeneity induces spurious negative duration dependence" - a mixed population of constant-risk individuals imitates a population that is getting safer.
Their headline finding is blunter still: "even when aggregate retention rates are monotonically increasing, the individual-level churn probabilities are unlikely to be declining over time."
The practical consequence: a falling hazard is not evidence that your product is getting stickier. It is equally consistent with a product that never gets stickier and a customer base that is sorting itself. Distinguishing the two requires segmenting by something you knew at signup - plan, segment, acquisition channel, ICP fit score - and checking whether the hazard still falls within each segment. If it flattens once you segment, you were watching sorting.
This also means the tempting inference from a flat retention plateau - that the plateau height equals your real market - is one reading among several. Sorting produces the same picture.
What this is not
Three neighbouring problems get confused with this one, and the distinctions matter because the fixes differ:
- Immortal time bias is a time-alignment error: adopters are credited with time during which they could not have churned. That is a bug in how you assign person-time. The hazard curve is about the shape of risk across tenure once alignment is correct.
- The healthy adherer effect is a selection error: the accounts that complete onboarding differ from those that do not, before the onboarding does anything. That is about who is in the treated group.
- Survivorship bias is about who you never reached at all.
A hazard curve can be perfectly computed and still be misread, which is what the sorting trap above describes. All four can be present in the same chart.
Which method belongs in which regime
Because the causes differ, the evidence has to be collected differently.
Infant-mortality churn is recent, so recall is good and the sample is large - this is where a structured exit study works. Useful-life churn is exogenous, so the useful question is not about your product; it is about what changed at the account, and much of it will be uncontrollable. Wear-out churn is the hardest, because the cause accumulated slowly and the customer often cannot name it. Nobody remembers the month the workarounds became intolerable. That regime needs longitudinal contact with still-active tenured accounts, not exit interviews - by the time they cancel, the reconstruction is post-hoc.
The awkward truth is that this is three recruiting screens, three discussion guides, and three analyses. Most teams run one, because three was never affordable.
How Koji changes the economics
Running a separate study per tenure regime is a scheduling and cost problem before it is a methods problem. Koji removes the constraint that makes teams collapse three studies into one:
- Segment-matched studies in parallel. Define three cohorts by tenure band and field all three at once with AI-moderated interviews. The marginal cost of the third study is close to the cost of the first, so the per-regime design stops competing with the deadline.
- Identical probing across bands. A human moderator running month-1 and month-24 interviews will probe the interesting one harder, and the resulting difference is inseparable from the method - the mechanism described in our guide to surveillance bias. An AI moderator applies the same brief and the same follow-up logic to every band, which is what makes cross-regime comparison legitimate.
- A detection-independent baseline. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no. The five closed types mean the same thing regardless of how long the conversation ran, so they give you a fixed yardstick across tenure bands. If open-ended themes diverge sharply between month 1 and month 24 but the scale and ranking data does not, you are looking at a probing artifact rather than a real regime difference.
- Continuous fielding for the wear-out regime. Wear-out is the regime that defeats exit research. A standing quarterly study of tenured active accounts catches the accumulation while the customer can still describe it.
Teams adopting AI-assisted research consistently report time-to-insight measured in days rather than the six-to-eight weeks a three-arm qualitative study traditionally consumes - which is the difference between segmenting your hazard curve and rounding it to one number.
A working checklist
- Plot the hazard, not just retention: churned in month t divided by accounts that reached month t.
- Plot it on a log scale if your month-1 spike compresses everything else flat.
- Mark the three regimes and check whether the tail turns back up. A rising tail is the expensive finding.
- Re-plot within segments you knew at signup. If the decline disappears, you were watching sorting, not loyalty.
- Do not average a churn rate across regimes for any decision that funds work.
- Assign one study per regime, with the recruiting screen written from the hazard chart.
- For the wear-out regime, interview active tenured accounts, not churned ones.
- Re-run the whole analysis quarterly - the regime boundaries move when onboarding or pricing changes.
Frequently asked questions
What is the difference between a retention curve and a hazard curve?
A retention curve is a survival function: the share of an original cohort still active at each point in time. It can only decrease. A hazard curve is conditional: among accounts that survived to month t, the share that left during month t. It can rise or fall. Two products with identical retention curves can have completely different hazard shapes, and the hazard is what tells you which intervention will pay.
How do I calculate a churn hazard rate?
For each tenure month, divide the number of accounts that churned during that month by the number of accounts that entered that month alive. Do not divide by the original cohort size - that gives the unconditional rate, which understates late-life risk because the denominator includes accounts that already left. The reliability-engineering worked example is .053 / .692 = .077.
Does a flattening retention curve prove product-market fit?
Not on its own. A flattening curve is exactly what you would see if every customer had a fixed churn probability and the high-risk ones simply left first. Fader and Hardie show that increasing cohort-level retention can be entirely a sorting effect from cross-sectional heterogeneity, with no individual becoming more loyal. Segment by a signup-time attribute and see whether the flattening survives.
How many tenure buckets should I use?
Enough that each bucket has a stable denominator - as a rule of thumb, at least a few hundred accounts entering the bucket before the rate stops being noise. Monthly buckets for the first quarter, where the action is, then quarterly buckets afterwards is a reasonable default for most B2B books.
Why does my churn hazard go up again for long-tenured accounts?
That is the wear-out regime, and it is the pattern worth investigating first because those accounts are usually your largest. Common causes are accumulated workarounds that finally exceed tolerance, outgrowing the data model or permission structure, and a competitor closing a gap that mattered at renewal. It is rarely a single event, which is why churned-customer interviews reconstruct it badly.
Can I use this for user-level retention, not just accounts?
Yes. The arithmetic is identical for users, seats, or any unit that can leave. The regimes tend to compress - user-level infant mortality often plays out over days rather than months - but the three-regime structure and the sorting trap both apply unchanged.
Related Resources
- Cohort Analysis: How to Read Retention and Find the Why - the survival-curve side of the same data
- Immortal Time Bias in Retention Analysis - the time-alignment error that fakes a feature effect
- The Healthy Adherer Effect - why onboarding completers always look better
- Structured Questions Guide - the six question types and when to use each
- Customer Retention Research - the full retention research program
- Case-Control Research for Churn and Lost Deals - designing the comparison properly
Related Articles
Case-Control Research: How to Study Churn and Lost Deals Without Fooling Yourself (2026)
Every churn interview and win-loss study is a case-control design, whether or not anyone says so. Epidemiology has spent seventy-five years learning how these studies go wrong - control selection, admission bias, recall bias, and base rates.
Cohort Analysis: How to Read Retention and Find the "Why" (2026)
Cohort analysis groups users by a shared starting point and tracks their behavior over time, revealing retention patterns that aggregate metrics hide. This guide explains how to build and read cohort tables, interpret the retention curve, and pair the numbers with qualitative research to explain them.
Customer Retention Research: The Complete 2026 Playbook for Reducing Churn Before It Happens
A practitioner's guide to customer retention research — how to combine churn interviews, stay interviews, NPS follow-ups, and continuous voice-of-customer programs to reduce churn 25% or more. Includes question templates, sampling frameworks, and how AI-moderated research scales retention listening across your entire customer base.
The Healthy Adherer Effect: Why Users Who Finish Onboarding Always Retain Better (2026)
Users who complete your onboarding checklist retain better. So do users who adhere to a placebo. The healthy adherer effect explains why adoption metrics overstate feature impact, why adjusting for covariates does not fix it, and what to do instead.
Immortal Time Bias: Why Feature Adopters Always Look More Loyal Than They Are (2026)
Immortal time bias makes every feature-adoption retention chart overstate the feature. Learn how the bias works, why product data is the worst case, and the three fixes.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Surveillance Bias: Why the Team That Measures Best Looks Worst (2026)
The harder you look, the more you find. Surveillance and lead time bias make well-instrumented teams look worse and useless interventions look effective. Here is how to tell the difference.
Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story
Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.