Back to docs
Research Methods

Surveillance Bias: Why the Team That Measures Best Looks Worst (2026)

The harder you look, the more you find. Surveillance and lead time bias make well-instrumented teams look worse and useless interventions look effective. Here is how to tell the difference.

Surveillance Bias: Why the Team That Measures Best Looks Worst

Answer first: Any metric that counts detected events measures three things multiplied together: how often the event happens, how hard you look for it, and whether anyone records it. Only the first is what you care about. When you compare teams, products or periods with different detection intensity, you are mostly measuring the looking. This is surveillance bias, and its close relative lead time bias means an intervention can improve every survival number you report while changing no outcome at all. The test that separates them is simple, and almost nobody runs it.

The finding that should end cross-team metric comparisons

In 2013, Karl Bilimoria and colleagues published a study in JAMA (310(14):1482-1489) that is the cleanest demonstration of this problem ever produced. They merged 2010 Hospital Compare and American Hospital Association data from 2,838 hospitals, then used 2009-2010 Medicare claims covering 954,926 surgical patient discharges from 2,786 hospitals across 11 major operations.

Postoperative venous thromboembolism, a blood clot, is a widely reported hospital quality metric that was slated for use in pay-for-performance programmes. The logic is intuitive: good hospitals prevent clots, so good hospitals should report fewer clots.

The data said the opposite.

Hospital structural quality quartileClot-prevention adherenceRisk-adjusted clot rate
Lowest quality quartile93.3 percent4.8 per 1,000
Highest quality quartile95.5 percent6.4 per 1,000

Hospitals with more accreditations and more quality initiatives did better on prevention and worse on the outcome metric, both differences significant at P < .001. Greater adherence to prevention was weakly associated with worse risk-adjusted event rates.

Then the authors found the cause. They sorted hospitals by how often they ordered the imaging that detects clots:

Imaging use quartileDiagnostic imaging rateRisk-adjusted clot rate
Lowest32 studies per 1,0005.0 per 1,000
Highest167 studies per 1,00013.5 per 1,000

A five-fold difference in looking produced a 2.7-fold difference in finding, rising in a clean stepwise fashion across quartiles (P < .001). The authors' conclusion was blunt: surveillance bias limits the usefulness of the measure both for hospitals trying to improve and for patients trying to choose one.

As Haut and Pronovost put it in their JAMA commentary on the phenomenon (305(23):2462-2463, 2011), the principle is that the greater the intensity of the search for a condition, the greater the likelihood it will be found.

Why this is your metrics dashboard

Substitute freely. The structure survives every substitution:

Your metricThe team that looks harderWhat the metric actually rewards
Production errors per releaseThe squad with the best error monitoring and alertingTurning off Sentry
Support tickets per 1,000 accountsThe team that made it easiest to file a ticketHiding the help widget
At-risk accounts flaggedThe CS org with the best health scoringScoring nothing
Accessibility defects foundThe team that actually runs an auditNot auditing
Security incidents reportedThe org with a functioning speak-up cultureDiscouraging reports
Usability problems per studyThe researcher who ran more sessions with tougher tasksRunning easier tasks

Every row is a case where the better-performing team posts the worse number, and where the cheapest way to improve the metric is to degrade the measurement. If your quarterly review compares detected-event counts across teams whose instrumentation differs, the review is ranking instrumentation.

The sign inversion: two biases pulling opposite ways

It is worth being precise about how this differs from immortal time bias, because the two are mirror images and teams routinely have both at once.

Immortal time biasSurveillance bias
What is unearnedSurvival timeDetection
Direction of errorFlatters the treatmentPenalises the better-measured group
Where it livesThe group definitionThe measurement process
Fixed byRealigning the clocksMeasuring or equalising the looking

One makes your feature look better than it is. The other makes your best team look worse than it is. Neither is visible in the metric itself, and no amount of statistical adjustment on the outcome fixes either one, because in both cases the corruption entered before the outcome was recorded.

Lead time bias: improving the number without improving anything

The second half of this problem is subtler and, in product work, more expensive.

The Mayo Lung Project was a randomised trial of lung cancer screening in 9,211 male smokers conducted between 1971 and 1983. The intervention arm was offered chest x-ray and sputum cytology every four months for six years. Marcus and colleagues published extended follow-up in the Journal of the National Cancer Institute (92(16):1308-1316, 2000), with a median follow-up of 20.5 years.

Here is what screening achieved.

Survival, measured from diagnosis, was clearly better in the screened arm (P = .0039). Among patients with resected early-stage disease, median survival was 16.0 years in the screened arm versus 5.0 years in the usual-care arm. On the metric that every hospital brochure reports, screening looked like a triumph: it more than tripled survival.

Mortality, the rate at which people actually died of lung cancer, was 4.4 per 1,000 person-years in the screened arm and 3.9 in the usual-care arm (P = .09). Screening did not reduce deaths. If anything the point estimate ran the wrong way.

Both numbers are correct. Survival is measured from the moment of diagnosis, and screening moves the moment of diagnosis earlier. If you find the disease three years sooner and the person dies on exactly the same day, survival from diagnosis rises by three years and nothing whatsoever has improved. That gap is the lead time.

A follow-up analysis by the same group (JNCI 98(11):748-756, 2006) tracked lung cancer incidence through 1999 and found 585 cases in the intervention arm against 500 in usual care. The excess persisted after 16 additional years of follow-up, which is the signature of overdiagnosis: finding disease that would never have caused a problem.

The three-way decomposition

The generalisable idea underneath all of this is worth stating as a formula, because it makes the failure mode obvious:

Observed event rate = true incidence x detection intensity x reporting propensity

You care about the first term. You control the second and third. Any change in your looking or your logging moves the observed rate without moving reality, and the observed rate is the only one you can see.

This yields three rules:

  1. Never compare detected-event rates across units with different instrumentation. Compare within a unit over a period when instrumentation did not change, or explicitly measure and adjust for the looking.
  2. When a detection improvement ships, expect the metric to get worse, and say so before it happens. Announce the expected step change in advance. A team that predicts its own metric regression retains credibility; a team that explains it afterwards does not.
  3. Report the denominator of looking alongside the numerator. Bilimoria's contribution was not discovering that clots vary. It was measuring the imaging rate. Publish "errors per release" next to "percentage of code paths instrumented" and the comparison becomes interpretable.

The one test that separates a real win from lead time

When someone claims an early-detection programme worked, ask a single question:

Did the number of bad outcomes fall, or did the time between detection and the bad outcome grow?

Only the first is a win. The second is arithmetic.

Concretely, if a customer success team says "accounts we flag survive 11 months from flag date, versus 4 months for unflagged accounts from when their problems started", they have reported a survival statistic and proved nothing. The correct measure is the churn rate of the whole eligible population before and after the programme, ideally with a randomised holdout of accounts that are scored but not contacted.

The holdout is what kills the argument, and it is also what almost nobody is willing to run, because it means deliberately not saving some accounts. That reluctance is exactly why so many retention programmes have been running for years with no evidence they do anything. The equivalent of the overdiagnosis finding is the account flagged as at-risk that was never going to churn: every one of those is a "save" your team gets credit for and a cost your business absorbed for nothing.

Surveillance bias inside qualitative research

This is not only a quantitative problem, and the version that hides inside interview studies is the one researchers most often miss.

A moderator runs 20 interviews over three weeks: 10 with enterprise customers and 10 with SMB customers. By week two, the enterprise conversations have surfaced something interesting, so the moderator probes those threads harder, asks more follow-ups, and spends longer on the topic. The report concludes that enterprise customers have more integration pain.

Do they? Or were they simply searched harder? Detection intensity differed systematically between the comparison groups, introduced by the researcher, and the finding is not separable from the method that produced it. Unequal probing depth across segments is surveillance bias with a human moderator as the imaging machine, and it appears in the limitations section of approximately no research reports.

The same mechanism operates across time. A researcher who runs a study before and after a redesign knows what changed and probes the changed areas harder in round two. More problems surface in the changed areas. That is not evidence the redesign made things worse.

The modern approach: how Koji helps

The fix for surveillance bias in research is a constant instrument, applied with the same intensity to every group and every round. That is difficult for humans, who get curious, tired, and biased toward what they already believe, and straightforward for a well-specified AI moderator.

  • Identical probing depth across segments. A Koji study runs the same brief for every participant. Enterprise and SMB respondents get the same questions with the same follow-up logic, so a difference between segments is a difference in what they said, not in how hard someone dug.
  • Structured questions give you a detection-independent baseline. The six types matter here specifically because closed types are immune to probing depth. A scale rating, a single_choice selection, a multiple_choice set, a ranking order and a yes_no answer mean the same thing regardless of how long the conversation ran. open_ended questions with AI follow-up are where depth varies, so pair them: if the open-ended themes diverge between segments but the scale and ranking data do not, you are looking at a probing artifact rather than a real difference. The structured questions guide covers how to build that pairing.
  • Consistency across rounds. Re-running the identical brief before and after a change means round two is not searched harder than round one. Manual longitudinal research almost never achieves this, because the moderator has learned things in between.
  • Automatic thematic analysis applies one standard. Every transcript is coded by the same process, so a theme is not more likely to be recorded in the segment the researcher found interesting.
  • Real-time reporting exposes the volume difference. You can see immediately whether one segment produced more content than another, which is the diagnostic signal for unequal detection.

Against legacy tooling this is a structural advantage rather than a speed one. A SurveyMonkey field applies a constant instrument but cannot probe at all, so it never gets the mechanism. A human interview programme probes well but cannot hold intensity constant across 40 conversations. AI moderation is the only approach that delivers both: adaptive follow-up within a fixed, auditable instruction set. Teams using AI-assisted research report substantially faster time-to-insight, but the more durable benefit here is that the instrument does not drift between the groups you are comparing.

A caveat worth stating plainly, because it is the honest version of this claim: adaptation in the probe is not the same as inconsistency in the instrument, provided the instruction set was fixed before fielding. What varies between two Koji interviews is the realisation of a rule, not the rule.

Frequently asked questions

What is surveillance bias?

Surveillance bias is the distortion that occurs when the intensity of searching for an event differs between the groups being compared. Because looking harder finds more, the more closely monitored group posts higher event rates regardless of true incidence. Bilimoria and colleagues demonstrated it across 954,926 surgical discharges: hospitals in the highest imaging quartile reported 13.5 clots per 1,000 versus 5.0 in the lowest.

What is the difference between surveillance bias and lead time bias?

Surveillance bias inflates event counts because you searched harder. Lead time bias inflates survival times because you detected the event earlier without changing the outcome. In the Mayo Lung Project, screening produced median survival of 16.0 years versus 5.0 years for resected early-stage disease while lung cancer mortality was 4.4 versus 3.9 per 1,000 person-years, a non-significant difference.

How do I know whether my early-warning programme actually works?

Ask whether the number of bad outcomes fell, or whether only the time between detection and the outcome grew. Measure the churn rate of the entire eligible population before and after, not survival from the flag date. The definitive design is a randomised holdout of accounts that are scored but not contacted.

Can I fix surveillance bias with statistical adjustment?

Not on the outcome alone. The corruption enters before the event is recorded, so adjusting the outcome cannot recover it. The workable fixes are to measure detection intensity explicitly and report it alongside the event rate, to compare only within units whose instrumentation did not change, or to standardise the looking across groups.

Does surveillance bias affect qualitative research?

Yes, and it is commonly missed. When a moderator probes one segment harder than another, more issues surface in the harder-probed segment, and the finding cannot be separated from the method. The defence is a constant instrument: the same brief, the same follow-up logic, and closed question types whose answers do not depend on conversation length.

Why did better hospitals score worse on the quality measure?

Because the measure counted detected clots. Higher-quality hospitals had better prevention adherence, 95.5 percent versus 93.3 percent, but also imaged far more, and imaging is what turns a clot into a recorded event. Their risk-adjusted rates were 6.4 versus 4.8 per 1,000. The metric ranked surveillance intensity while appearing to rank quality.

Related Resources

Related Articles

Blind Analysis: How to Analyze Research Before You Know the Answer

Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.

The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment (2026)

Most of what matters in product research cannot be randomised. Bradford Hill nine viewpoints are the framework for building a defensible causal case without an A/B test.

Correlation vs. Causation: Why Your Metrics Lie (and How to Find the Real Why)

A practical guide to correlation versus causation for product and research teams: why the two get confused, the classic traps, how to establish real causation, and how qualitative interviews reveal the mechanism behind the numbers.

Immortal Time Bias: Why Feature Adopters Always Look More Loyal Than They Are (2026)

Immortal time bias makes every feature-adoption retention chart overstate the feature. Learn how the bias works, why product data is the worst case, and the three fixes.

Panel Conditioning: Why Your Most Reliable Participants Give You the Least Reliable Data (2026)

Panel conditioning is the measurement error you create by asking the same people again. Government statistical agencies have measured it for seventy years and it moves headline numbers by a full percentage point. Here is how to detect it in a product research panel and design around it.

Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)

Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.