Back to docs
Research Methods

Competing Risks: Why Your Retention Curve Overstates the Churn You Care About (2026)

Your retention curve treats acquisitions, downgrades and payment failures as if those accounts were still at risk of cancelling. That inflates the number. Here is the correction, the size of the error, and the interview that produces the missing field.

Answer first: if you compute a retention curve for one kind of churn and treat every other kind of exit as if the customer simply left the study, the number you get is too high - and it is too high even when the exits are unrelated to each other. The fix is not a better model. It is admitting that "churn" is four or five different events, estimating each one separately, and reporting the version of the curve that answers the question the room is actually asking.

Almost every retention chart in every board deck is a survival curve. Take a cohort, follow it forward, plot the share still alive at each month. The chart is correct as far as it goes. The problem is what it does with the accounts that leave for a reason you were not measuring.

A B2B account can exit your book in at least five distinct ways: it cancels outright, it downgrades to a free tier and stops being revenue, it gets acquired and absorbed into a parent contract, its champion leaves and the renewal quietly lapses, or its payment method fails and nobody notices. If your question is "how many accounts cancel," the other four are competing risks - events whose occurrence makes the event you care about impossible to observe. An account absorbed into a parent contract in month 7 cannot cancel in month 14. It is not "still at risk." It is gone.

The standard survival estimator does not know that. It treats the acquired account as censored - as if it merely walked out of your field of view while remaining fully capable of cancelling later. That assumption inflates your cancellation number, and the inflation is not small.

The five exits your retention curve treats as one

Exit typeWhat the account didWhat the standard curve assumesWhy that is wrong
Voluntary cancellationActively cancelled the contractIt cancelled - correctThis is the event you are estimating
Downgrade to freeStopped paying, kept usingStill at risk of cancellingRevenue already gone; the risk profile is now completely different
Acquisition or mergerAbsorbed into a parent contractStill at risk of cancellingIt can never cancel; the decision maker no longer exists
Champion departure lapseRenewal quietly not actionedStill at risk of cancellingThe mechanism is organisational, not product dissatisfaction
Payment failureCard expired, dunning failedStill at risk of cancellingInvoluntary; often reversible within days

Each row that is not the first one is a competing risk for the first one. And each of them is also its own event with its own curve, its own drivers, and its own fix. Lumping them produces a number that describes no decision you can make.

What censoring actually assumes, in plain English

Censoring is the mechanism that lets survival analysis handle incomplete follow-up. If an account signed up four months ago and your window is twelve months, you censor it at month 4: you know it survived four months, you do not know what happens after, and the estimator correctly stops counting it.

That works because the reason for the censoring - your data collection ended - has nothing to do with whether the account was about to churn. Administrative censoring is uninformative. Competing events are not. When you censor an acquired account, you are telling the estimator: assume this account would have gone on to cancel at the same rate as the accounts still on your books. It cannot. It does not exist as an independent buyer anymore.

The consequence is stated bluntly in the standard methodological reference for this problem. Peter Austin, Douglas Lee and Jason Fine, writing in Circulation (133(6):601-609, 2016), put it this way: the use of the Kaplan-Meier survival function to estimate incidence "results in estimates of incidence that are biased upward, regardless of whether the competing events are independent of one another."

That last clause is the one that surprises people. The usual defence - "acquisitions are random with respect to product satisfaction, so it washes out" - does not rescue you. Independence is not the issue. The issue is that a customer who has already exited by another door is being counted as someone still standing in front of your door.

The size of the error, measured

Austin, Lee and Fine worked their example on the Enhanced Feedback for Effective Cardiac Treatment (EFFECT) heart failure cohort: 16,237 patients across 103 Ontario hospitals, followed for five years, with two mutually exclusive outcomes - cardiovascular death and non-cardiovascular death.

  • Estimated using the complement of the Kaplan-Meier survival function, five-year incidence of cardiovascular death came out at 43.0 percent.
  • Estimated using the cumulative incidence function (CIF), which accounts for the competing event, the same quantity came out at 36.8 percent.
  • The naive estimate was 6.2 percentage points too high - a relative overstatement of about 17 percent of the true figure.

There is a second tell, and it is one you can check on your own data this afternoon without any new software. When you estimate each exit type separately with the naive method and add the results, the pieces sum to more than the whole. Austin and colleagues observed exactly this: the sum of the two Kaplan-Meier incidence estimates exceeded the Kaplan-Meier estimate of all-cause mortality. If your cancellation curve plus your downgrade curve plus your acquisition curve add up to more than your total attrition curve, you have found the bug. Probabilities of mutually exclusive events cannot sum past the probability of any of them happening.

How common is the mistake? Koller, Raatz, Steyerberg and Wolbers reviewed 50 clinical studies of populations obviously subject to competing risks, all published in high-impact journals, and found competing risks issues in 70 percent of the articles; only two studies explicitly used the cumulative incidence function in place of biased Kaplan-Meier estimates (Statistics in Medicine 31(11):1089-1097, 2012). This is not an exotic failure. It is the default in a field with statistical reviewers. Your growth dashboard does not have statistical reviewers.

Two questions, two curves - and you must pick one on purpose

The reason this cannot be solved by "just use the right formula" is that there are two legitimate formulas, and they answer different questions.

You want to knowUseWhat it tells youTypical audience
What drives the rate of cancelling among accounts still on the booksCause-specific hazardThe instantaneous rate of cancellation among accounts that have not yet exited by any routeProduct and research: this is the etiologic question
What share of this cohort will actually cancel by month 12, in the real world where other exits happenCumulative incidence function / subdistribution hazard (Fine-Gray)Absolute risk of this specific exit, in the presence of everything else that can happenFinance, forecasting, and anyone sizing a save programme

Austin, Lee and Fine are explicit about the division of labour: the cause-specific model "may be better suited for addressing etiologic questions, whereas the latter model may be better suited for estimating a patient's clinical prognosis." Translated: use the cause-specific hazard to work out why accounts cancel; use cumulative incidence to work out how many will. A finding stated with the wrong one of these is not a rounding error, it is an answer to a question nobody asked.

The formal machinery for the second one is Fine and Gray's subdistribution hazard model (Journal of the American Statistical Association 94:496-509, 1999). You do not need to fit it to benefit from this article - plotting cumulative incidence functions instead of one-minus-Kaplan-Meier is the whole of the practical win - but if you have a churn risk model, the distinction is not academic.

The risk model that flags twice as many accounts as it should

Here is where the statistics turns into a wasted quarter of customer success capacity. Wolbers and colleagues compared a standard Cox survival model against a Fine-Gray model for predicting coronary heart disease in older women, a population with heavy competing mortality. The standard model classified 18 percent of subjects as high risk. The competing-risks model classified 8 percent. Same data, same outcome, same threshold - more than twice as many people flagged by the model that ignored the competing events.

Now map that onto a churn-risk score. If your model was trained on a target of "churned in the next 90 days" where acquisitions, downgrades and dunning failures were folded in or censored, your high-risk list is systematically too long. Every extra name on that list is a save play, an executive email, a discount authorisation, and an hour of a CSM's week spent on an account that was never going to cancel - or on one that was going to be acquired no matter what you did.

What this is not

Three neighbouring ideas that people confuse with this one, and where each of them actually lives:

  • This is not the hazard curve. The churn hazard curve is about when risk is concentrated across tenure - the three regimes, and the trap where a falling hazard is sorting rather than loyalty. Competing risks is about which exit you are measuring. You can and should do both: a cause-specific hazard curve, plotted per exit type, is the fully specified version of that chart.
  • This is not the churn reason taxonomy. Churn surveys already sort cancellations into stated reasons - price, missing feature, switched vendor. That is a categorisation of the cancellation event. Competing risks is about the events that are not cancellations at all and never generate a churn survey response, because nobody fills in a cancellation form when their company gets acquired.
  • This is not survivorship bias. Survivorship bias is about which customers you talk to. This is about the arithmetic of a curve you compute from all of them.

The research move: the exit type is a question, not a database field

The uncomfortable part of implementing this is that your billing system does not know why an account left. It knows the subscription ended. Whether that was a deliberate cancellation, an absorbed contract, a lapsed renewal nobody actioned, or a card that failed while the champion was on parental leave is a fact about a human decision, and it lives in exactly one place: the humans.

This is the strongest practical argument for talking to churned accounts that has ever been available, because it is not a soft argument about empathy. Misclassifying exits does not just lose you a reason code. It biases the denominator of every retention number you report. You cannot fix that in the warehouse.

Three classification questions that separate the exits cleanly, and that almost nobody asks:

  1. "At the point the subscription ended, was there a person whose job it was to decide whether to renew?" (Separates a genuine decision from an organisational lapse.)
  2. "Did your company's structure or ownership change during the contract term?" (Catches acquisitions and consolidations, which almost never make it into a reason code.)
  3. "If the payment had gone through automatically, would you still be using it today?" (Separates involuntary churn that was truly involuntary from involuntary churn that was a decision someone was relieved not to have to make.)

That third one matters more than teams expect. Payment-failure churn is routinely written off as an ops problem, and some of it genuinely is. But a customer who had already decided to stop and simply let the card lapse is a voluntary cancellation wearing an involuntary costume, and it belongs in a different curve.

How Koji helps

The classification work above is why competing-risks analysis usually never happens: it needs a conversation with every exiting account, and nobody has the interviewer hours. Twenty exits a month at 30 minutes each, plus scheduling, plus transcription, plus coding, is most of a full-time role. So teams settle for a dropdown in the cancel flow, which the acquired accounts never see.

Koji removes the constraint by making the exit interview automatic and simultaneous rather than scheduled and serial:

  • AI-moderated interviews run the moment an account exits, at any volume, in parallel. There is no queue, so the acquired account and the dunning-failure account both get interviewed, not just the ones who clicked through a cancel flow.
  • Structured questions give you the classification as typed data rather than prose you have to code by hand. Koji supports six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and a competing-risks classifier is a single single_choice exit-type question plus a yes_no on the ownership change, with an open_ended follow-up where the AI probes whatever the first two answers imply. That single choice field is the column your survival analysis has been missing.
  • Automatic thematic analysis then runs within each exit type, which is the part that changes the roadmap. The themes behind champion-departure lapses have almost nothing in common with the themes behind price cancellations, and pooling them produces a theme list that is an average of two unrelated populations.
  • Voice interviews raise completion among exiting accounts who will not type but will talk for four minutes, which matters because the exits you are least likely to hear from are exactly the non-cancellation ones.
  • Real-time reporting means the exit-type distribution updates as responses land, so you can watch whether the mix is shifting a quarter before the blended number moves.

The comparison with legacy tooling is stark. A traditional survey tool can capture a cancellation reason from someone who reaches the cancel screen. It cannot interview the 40 percent of exits that never reach a cancel screen, because those accounts did not cancel - they were absorbed, lapsed, or dunned out. Getting to them requires outbound, moderated conversation at a volume no team staffs for manually.

Two things that depend on getting this right

Exit-type classification is not a tidiness exercise. Two downstream numbers inherit its errors directly.

The first is your cost per save. Number needed to treat - how many accounts you must reach to retain one - is one divided by the absolute risk reduction, and the absolute risk reduction is the baseline risk multiplied by whatever relative improvement your intervention delivers. An inflated baseline makes the number needed to treat look smaller, which makes every save programme costed against it look cheaper than it is.

The second is which intervention you compare it against. A large share of the exits you classify as voluntary cancellation trace back not to a missing feature but to an unresolved problem that stayed unresolved for weeks. That is a time-to-repair problem, not a prevention problem, and it is usually the cheaper of the two to fix. You cannot see it at all while every exit is one undifferentiated event.

A working checklist

  1. Enumerate every way an account can leave your book. If your list has one item, that is the finding.
  2. Add an exit-type field to the churn table, and populate it from interviews, not from inference.
  3. Plot cumulative incidence functions per exit type, not one-minus-Kaplan-Meier.
  4. Run the sum check: the per-type incidences must not exceed all-cause attrition. If they do, you are still using the naive estimator.
  5. Report cause-specific hazards when you are explaining why, and cumulative incidence when you are forecasting how many. Label which one is on the chart.
  6. Re-check any churn risk model whose training target was "churned" without an exit type. Expect the high-risk list to shrink materially.
  7. Re-run the classification quarterly. The mix shifts with your market - a consolidation wave changes your acquisition-exit share without changing your product at all.

Frequently asked questions

What is a competing risk in churn analysis?

A competing risk is any exit that makes the exit you are measuring impossible to observe. If you are estimating voluntary cancellation, then acquisition, downgrade to a free tier, an unactioned renewal, and payment-failure churn are all competing risks. The defining property is that the account cannot subsequently experience your event of interest - an acquired company cannot later cancel, because the buying entity no longer exists.

Why does Kaplan-Meier overestimate churn when competing risks exist?

Because it treats competing events as censoring, and censoring assumes the account is still at risk. Austin, Lee and Fine (Circulation, 2016) show this bias is upward "regardless of whether the competing events are independent of one another." In their worked example, the naive estimate of five-year cardiovascular death was 43.0 percent against a correct cumulative incidence of 36.8 percent - 6.2 percentage points too high.

What should I use instead of one minus Kaplan-Meier?

The cumulative incidence function, which estimates the absolute probability of a specific exit type occurring by a given time in the presence of the other exits. For regression, the cause-specific hazard model answers questions about what drives the rate of an exit among accounts still at risk, and the Fine-Gray subdistribution hazard model (JASA, 1999) answers questions about absolute risk. Pick based on whether you are explaining or forecasting.

How do I tell whether my retention curve has this problem?

Estimate each exit type separately with your current method and add them up. If the sum exceeds your all-cause attrition curve, the estimator is inflating each piece - mutually exclusive probabilities cannot sum past the total. This check takes one query and no new tooling.

Is involuntary churn a competing risk or the same event?

Treat it as its own event until you have evidence otherwise. Payment-failure churn has a different mechanism, a different fix, and often a different time signature than voluntary cancellation. Some of it is genuinely a billing problem and some of it is a decision in disguise, which is why the interview question "if the payment had gone through automatically, would you still be using it today?" is worth asking of every dunning exit.

Does this change how many interviews I need?

It changes who you interview more than how many. You need coverage of each exit type, not just of the cancel-flow population, because the non-cancellation exits are precisely the ones no survey reaches. In practice that means outbound interviews to accounts that never touched a cancel screen - which is why an AI-moderated approach that runs in parallel is what makes the analysis feasible at all.

Related Resources

Related Articles

Case-Control Research: How to Study Churn and Lost Deals Without Fooling Yourself (2026)

Every churn interview and win-loss study is a case-control design, whether or not anyone says so. Epidemiology has spent seventy-five years learning how these studies go wrong - control selection, admission bias, recall bias, and base rates.

The Churn Hazard Curve: Why One Churn Rate Hides Three Different Problems (2026)

Your monthly churn rate averages three unrelated problems into one number. Learn to plot the churn hazard by tenure, read the three regimes, and avoid the sorting trap that makes a flattening curve look like product-market fit.

How to Build Churn Surveys That Actually Save Customers

Learn how to design churn surveys that uncover real cancellation reasons, optimize exit flows, and feed win-back strategies. Use AI conversations to empathetically engage departing customers.

Cohort Analysis: How to Read Retention and Find the "Why" (2026)

Cohort analysis groups users by a shared starting point and tracks their behavior over time, revealing retention patterns that aggregate metrics hide. This guide explains how to build and read cohort tables, interpret the retention curve, and pair the numbers with qualitative research to explain them.

Average Customer Lifetime Is Not 1 Divided by Your Churn Rate (2026)

The LTV = ARPU / churn formula assumes a constant hazard. Three defensible methods on the same book of business give 9, 20, and 67 months. What to report instead, and why reliability engineering solved this sixty years ago.

Availability, Not Uptime: Why Time-to-Repair Is Half Your Retention Equation (2026)

Availability is a ratio with two terms, and product teams fund only one of them. Halving repair time and halving failure rate produce exactly the same result. Here is the arithmetic, the invisible parts of the customer repair clock, and how to measure them.

Number Needed to Treat: How Many Users You Must Reach to Keep One (2026)

Every effect in your deck is a rate. None of them is a count of people. Number needed to treat converts a percentage lift into the only figure a roadmap can cost, and the evidence says the persuasive format is the misleading one.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.