Back to docs
Research Methods

Immortal Time Bias: Why Feature Adopters Always Look More Loyal Than They Are (2026)

Immortal time bias makes every feature-adoption retention chart overstate the feature. Learn how the bias works, why product data is the worst case, and the three fixes.

Immortal Time Bias: Why Feature Adopters Always Look More Loyal Than They Are

Answer first: If you compare retention between users who adopted a feature and users who did not, and you start both clocks at signup, the result is wrong before you look at it. Everyone in the adopter group had to survive long enough to adopt. That stretch of guaranteed survival, which epidemiologists call immortal time, gets silently credited to the feature. The bias always runs in favour of the thing you are evaluating, it grows in proportion to the length of the adoption window, and it is worst when most churn happens early, which is precisely the shape of every SaaS retention curve. The fix is a landmark analysis or a time-varying exposure model, and neither works unless you recorded when adoption happened.

The chart that launched a thousand bad roadmaps

Someone on your team pulls a cohort, splits it by whether the account used the integrations feature in its first 90 days, and plots 12-month retention. Adopters retain at 71 percent. Non-adopters retain at 38 percent. The slide says: drive integration adoption.

The chart is not measuring what the slide claims. To be counted as an adopter, an account had to still be alive on the day it adopted. An account that churned in week two never had the chance to be classified as an adopter, so it lands in the comparison group by construction. The adopter group has been handed a block of survival time that no feature caused, and the non-adopter group has been handed every early churner in the business.

This is not a subtle statistical quibble. It is the difference between a real effect and no effect at all.

What immortal time actually means

The canonical definition comes from Samy Suissa, whose 2008 paper in the American Journal of Epidemiology (167(4):492-499) is the standard reference: immortal time is "a span of cohort follow-up during which, because of exposure definition, the outcome under study could not occur."

Read that phrase carefully: because of exposure definition. The immortality is not a property of the people. It is a property of how you defined the groups. You wrote a rule that says "adopters are people who did X", and that rule quietly requires survival up to the moment of X.

Suissa traces the problem back to the 1970s, when cohort studies of heart transplantation reported dramatic survival benefits. Patients had to live long enough to receive a transplant; time spent waiting on the list was credited to the transplanted group. The benefit shrank sharply once the time was allocated correctly.

The proof that you can manufacture an effect from nothing

The most damning demonstration is Suissa's own 2007 study in Pharmacoepidemiology and Drug Safety (16(3):241-249). He took a cohort of 3,315 patients with chronic obstructive pulmonary disease from the Saskatchewan Health databases, all hospitalised for cardiovascular disease, and followed them for up to a year.

He then tested two drug classes that have no known cardiovascular benefit whatsoever: gastrointestinal drugs and inhaled beta-agonists. Using the flawed cohort design, both looked protective against death:

Drug class (no known cardiac benefit)Biased analysis, rate ratio for deathCorrect person-time analysis
Inhaled beta-agonists0.73 (95% CI 0.57 to 0.93)0.98 (95% CI 0.77 to 1.25)
Gastrointestinal drugs0.78 (95% CI 0.61 to 0.99)0.94 (95% CI 0.73 to 1.20)

Both biased results were statistically significant. Both correct results were flat. The design alone produced an apparent 22 to 27 percent reduction in mortality from drugs that do nothing for the heart.

Translate that into your world: the analysis method is capable of making a feature that does nothing look like it cuts churn by a quarter, with a confidence interval that excludes zero. If your integrations chart shows a 20 percent retention lift, you cannot tell from the chart whether it is a real effect or an artifact of the same size.

The Oscar study, and why it is the perfect teaching case

In 2001, Redelmeier and Singh published a widely covered study in Annals of Internal Medicine reporting that Academy Award-winning actors and actresses lived almost four years longer than their less successful peers. The finding travelled the world as evidence that status and social rank affect longevity.

In 2006, Sylvestre, Huszti and Hanley reanalysed the same data in the same journal (Annals of Internal Medicine 145(5):361-363). Their diagnosis, in their own words, was that the original method "credited an Oscar winner is years of life before winning toward survival subsequent to winning." Once they used methods that avoided immortal time bias, the survival advantage fell to roughly one year and was no longer statistically significant.

To win an Oscar you must be alive at the ceremony. To adopt a feature you must be alive at adoption. The structures are identical. The Oscar study is your feature-adoption chart wearing a tuxedo.

Why product data is the worst possible case

Here is the part that should genuinely worry anyone running growth analytics, and it comes straight out of Suissa's 2008 quantification of the bias.

Two findings matter. First, the bias in the rate ratio increases in proportion to the duration of immortal time. A 90-day adoption window carries roughly three times the distortion of a 30-day window. Second, and more important, the bias is more pronounced when the hazard function is decreasing rather than constant. Suissa demonstrates this by comparing a Weibull distribution with a decreasing hazard against an exponential distribution with a constant hazard.

A decreasing hazard means the risk of the event is highest at the start and falls over time. That is the exact shape of a product retention curve. Most churn happens in the first weeks; users who survive onboarding churn at a much lower rate thereafter. Product analytics operates permanently in the regime where this bias is at its most severe, and it typically uses adoption windows measured in months.

So the two conditions that make immortal time bias worst, a long exposure window and a steeply decreasing hazard, are both the default in SaaS retention analysis. This is not an edge case for your team. It is the base case.

How common is the mistake?

Very. A systematic review of systematic reviews presented at the Peer Review Congress by Shin, Kim, Yon, Lee, Rahmati, Solmi, Carvalho, Koyanagi, Smith and John P. A. Ioannidis searched PubMed/MEDLINE, Embase and the Cochrane Database through 31 July 2024. Across 182 studies covering 25 topics, 44.0 percent (80 studies) were affected by immortal time bias. Among the 21 topics where both affected and unaffected studies existed, 57.1 percent (12 of 21) showed discordant results between the two groups, and in 23.8 percent (5 of 21) there was outright evidence reversal: the pooled conclusion flipped between significant and non-significant once biased studies were removed.

That is the published, peer-reviewed medical literature, written by people with formal epidemiological training and subjected to review. There is no reason to believe a product analytics dashboard assembled in an afternoon does better.

The three cohort definitions that create it

Suissa identifies three ways the bias enters. All three have exact product analogues:

Cohort definitionHow it appears in medicineHow it appears in your product
Time-basedExposure defined by a prescription filled within a fixed window after entry"Adopted the feature within 90 days of signup"
Event-basedExposure defined by an event that must occur to qualify"Completed onboarding", "invited a teammate", "connected an integration"
Exposure-basedFollow-up starts at cohort entry but exposure is assigned from a later actFollow-up runs from signup, but the group label comes from behaviour at any later point

If your group definition contains the words "ever", "within the first", or "at any point", you almost certainly have immortal time in the data.

How to spot it in five minutes

Run this checklist against any retention comparison before you act on it:

  1. Does group membership require an action? If yes, that action takes time, and that time is immortal.
  2. Do both clocks start at the same event? If both groups are measured from signup but one group is defined by a later act, the design is broken.
  3. Could a user churn before being classifiable? If early churners can only ever land in one group, the comparison is rigged.
  4. How long is the window? Longer window, larger bias, proportionally.
  5. Where does churn concentrate? Front-loaded churn amplifies the bias further.
  6. Does the effect shrink as you shorten the window? A real effect should be reasonably stable. An artifact shrinks with the immortal time it feeds on.

Point six is the cheapest diagnostic available. Rerun the same chart with a 14-day, 30-day and 90-day adoption window. If the "effect" tracks the window length, you have measured the window, not the feature.

The three fixes

FixWhat you doWhen to use it
Landmark analysisPick a fixed landmark time, such as day 30. Classify every account by whether it had adopted by that point. Drop anyone who churned before the landmark. Start the survival clock at the landmark for both groups.The default. Simple, transparent, explainable to a stakeholder in one sentence.
Time-varying exposureTreat adoption as a state that changes. Each account contributes non-adopter person-time until adoption and adopter person-time afterwards.When you have precise adoption timestamps and want to use all the data.
New-user designRestrict to accounts observed from their true start, and define exposure at a point that cannot depend on survival.When you control instrumentation and can design the measurement before the fact.

Landmark analysis is the one to reach for first. It costs one line of filtering logic, it discards the immortal time rather than misallocating it, and it produces a number you can defend in a roadmap review. Its price is that you throw away accounts that churned before the landmark, and you must say so out loud.

What no analysis can fix

Every one of these fixes needs one thing: an accurate record of when exposure happened. If your event log records that an account has the integration enabled but not the date it was enabled, no statistical method will recover the sequence. Time-alignment is an instrumentation decision made before the data exists, in the same way that a missing control group is a recruiting decision, not an analysis decision.

And even a correctly aligned chart tells you nothing about mechanism. Suppose the landmark analysis survives and adopters really do retain nine points better. You still do not know whether the integration causes retention or whether the kind of team that connects an integration in week one was always going to stick around. That is confounding by indication, and it is invisible to every method above. The only way to distinguish them is to ask the people involved what was happening at the time, and in what order.

This is why time-bias correction and qualitative research are complements rather than substitutes. The statistics tell you whether the association survives correct time alignment. The interviews tell you whether there is a mechanism worth building a roadmap on. For the broader framework covering how to move from a corrected association to a defensible causal claim, see the Bradford Hill criteria for product research.

The modern approach: how Koji helps

Correcting immortal time bias is arithmetic. Establishing the sequence and the mechanism is research, and that is the expensive half. Traditionally it means recruiting adopters and non-adopters, scheduling calls over three weeks, moderating each one, transcribing, and coding the transcripts by hand: a multi-week project that most teams simply skip, which is exactly why the uncorrected chart ends up driving the roadmap.

Koji collapses that loop:

  • AI-moderated interviews run in parallel, not in series. You field to adopters and non-adopters simultaneously and get results in days rather than weeks. The whole reason teams accept a biased chart is that the qualitative check felt too slow to bother with.
  • Structured questions pin down the sequence. Koji supports six question types, and the mix is what makes a timeline reconstructable. Use single_choice to establish which came first, the integration or the expansion decision. Use scale to measure how central the feature was to the renewal. Use ranking to force a comparison against the other things that happened that quarter. Use yes_no for clean eligibility gates, multiple_choice for the set of triggers, and open_ended with AI follow-up probing for the story behind the sequence. See the structured questions guide for how to combine them.
  • Automatic thematic analysis surfaces the mechanisms across every transcript at once, so "adopters were already planning to expand" shows up as a named theme rather than something a researcher happens to remember from call four.
  • Voice interviews capture the narrative detail, including hesitation and reordering, that a survey field never gets. People reconstruct timelines badly in text boxes and well in conversation.
  • A consistent instrument across both groups. A human moderator interviewing adopters and non-adopters over three weeks probes harder on whichever group is more interesting. That is a measurement difference between your comparison groups, introduced by the researcher. An AI moderator running the same brief asks both groups the same things with the same depth.

The last point matters more than it sounds, because unequal probing between groups is its own bias with its own name. It is covered in full in surveillance bias.

Unlike legacy survey tools such as SurveyMonkey, which capture a static snapshot and leave you to reconstruct chronology from a grid of answers, an AI-native platform can follow up in the moment: "You said you connected the integration after the renewal conversation. Walk me through what happened between those two." That is the question that resolves an immortal time problem, and no static form can ask it.

A worked workflow

  1. Pull the retention comparison as your team currently builds it. Record the headline number.
  2. Rerun it at 14, 30 and 90-day windows. If the effect scales with the window, flag the original as an artifact.
  3. Rebuild it as a landmark analysis at day 30. Record the corrected number and the count of accounts dropped.
  4. Take the corrected number to the team as the real estimate, with the drop count stated.
  5. Field a Koji study to both groups defined at the landmark, using structured questions to establish sequence and open-ended probing for mechanism.
  6. Report three things together: the corrected association, the mechanism, and the alternative explanations you could not rule out.

Teams that pair corrected quantitative estimates with fast qualitative mechanism checks stop shipping roadmap items justified by artifacts. That is the entire return on this method.

Frequently asked questions

What is immortal time bias in plain language?

It is the error you get when the way you defined your groups requires people to survive for a while before they can be counted in one of them. That guaranteed survival time gets credited to whatever the group represents. Samy Suissa's formal definition is that immortal time is a span of follow-up during which, because of how exposure was defined, the outcome could not occur.

Does immortal time bias always make the treatment look better?

Effectively always, in the standard setup. The exposed group is handed survival time it did not earn, and the unexposed group absorbs everyone who failed early. Suissa's demonstration is the clearest evidence: drugs with no cardiac benefit produced rate ratios of 0.73 and 0.78 under the biased design and 0.98 and 0.94 once person-time was allocated correctly.

How do I fix immortal time bias in a retention analysis?

Use a landmark analysis. Pick a fixed point such as day 30, classify accounts by whether they had adopted by then, exclude accounts that churned before the landmark, and start the survival clock at the landmark for both groups. If you have exact adoption timestamps, a time-varying exposure model uses more of the data. Both require knowing when adoption occurred.

Is this the same thing as survivorship bias?

No, and the distinction is worth keeping straight. Survivorship bias is about which people you reach at all: your in-app survey never reaches the users who already left. Immortal time bias affects people you did reach and did record, and comes from misaligning their clocks. You can have a perfectly representative sample and still have severe immortal time bias.

How big can the bias get in product data?

Larger than in most medical settings, because the two things that amplify it are both standard in SaaS. The bias grows in proportion to the length of the exposure window, and it is worse when the hazard is decreasing rather than constant. Retention curves have steeply decreasing hazards and adoption windows are usually 30 to 90 days, so product analytics sits in the worst corner of both.

Can I just add controls or adjust for confounders instead?

No. Immortal time bias is not confounding and adjustment does not touch it. It is a misallocation of follow-up time built into the group definitions, so it survives any amount of covariate adjustment. It has to be fixed in how time is assigned, which is why landmark and time-varying methods are the answer.

Related Resources

Related Articles

The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment (2026)

Most of what matters in product research cannot be randomised. Bradford Hill nine viewpoints are the framework for building a defensible causal case without an A/B test.

Case-Control Research: How to Study Churn and Lost Deals Without Fooling Yourself (2026)

Every churn interview and win-loss study is a case-control design, whether or not anyone says so. Epidemiology has spent seventy-five years learning how these studies go wrong - control selection, admission bias, recall bias, and base rates.

Cohort Analysis: How to Read Retention and Find the "Why" (2026)

Cohort analysis groups users by a shared starting point and tracks their behavior over time, revealing retention patterns that aggregate metrics hide. This guide explains how to build and read cohort tables, interpret the retention curve, and pair the numbers with qualitative research to explain them.

Correlation vs. Causation: Why Your Metrics Lie (and How to Find the Real Why)

A practical guide to correlation versus causation for product and research teams: why the two get confused, the classic traps, how to establish real causation, and how qualitative interviews reveal the mechanism behind the numbers.

Longitudinal Research: How to Track User Behavior and Attitudes Over Time

Longitudinal research captures how users change over time — not just a snapshot. This guide explains panel studies, cohort studies, and how AI-moderated interviews make multi-wave research feasible for any team.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Surveillance Bias: Why the Team That Measures Best Looks Worst (2026)

The harder you look, the more you find. Surveillance and lead time bias make well-instrumented teams look worse and useless interventions look effective. Here is how to tell the difference.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.