Back to docs
Research Methods

The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment (2026)

Most of what matters in product research cannot be randomised. Bradford Hill nine viewpoints are the framework for building a defensible causal case without an A/B test.

The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment

Answer first: The standard advice for establishing causation is to run a randomised experiment. It is correct advice, and it is unavailable for most of the questions product teams actually care about. You cannot randomly assign which customers churn, which companies are enterprises, which users hit the outage, or which buyer holds the budget. In 1965 Austin Bradford Hill set out nine viewpoints for weighing causation from observational evidence, precisely because the exposure in question, smoking, could never be randomised. Only one of the nine is experiment. This is the framework for the other cases, and it is deliberately not a checklist.

When the gold standard is not on the menu

Randomised experimentation deserves its reputation. Where you can run an A/B test, run one: random assignment neutralises confounders you have not even thought of, which is something no amount of analysis can do. That case is covered in correlation vs causation, and nothing here contradicts it.

But make an honest list of the causal questions your team asked last quarter:

  • Does poor onboarding cause churn?
  • Did the outage cost us the renewal?
  • Do teams that adopt our integration expand faster because of the integration?
  • Does pricing confusion cause abandoned trials?
  • Did the competitor's launch cause our win rate to fall?

Not one of these can be randomised. You cannot assign customers to have a bad onboarding, or to experience an outage, or to be in a market where a competitor launched. Some are impossible, some are unethical, and some are merely commercially unthinkable. The experiment is off the table, the question still needs an answer, and the decision gets made either way.

The choice is not between experimental evidence and observational evidence. It is between disciplined observational evidence and somebody's confident hunch.

Hill in 1965, and what he actually said

Sir Austin Bradford Hill delivered his inaugural presidential address to the Section of Occupational Medicine of the Royal Society of Medicine in 1965, published as "The Environment and Disease: Association or Causation?" in the Proceedings of the Royal Society of Medicine (58(5):295-300). Hill had already co-authored, with Richard Doll, the case-control work that established the link between smoking and lung cancer, and he had done it without a single randomised trial because no such trial was possible.

He proposed nine viewpoints, in this order: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. In his own words: "Here, then, are nine different viewpoints from all of which we should study association before we cry causation."

Two things about this are routinely got wrong.

First, Hill never called them criteria. He called them viewpoints. The word "criteria" was applied by later writers and it changed the meaning: criteria sound like a standard you either meet or fail, whereas viewpoints are angles from which to look.

Second, and this is the sentence that most people citing Hill have never read: "None of my nine viewpoints can bring indisputable evidence for or against the cause-and-effect hypothesis and none can be required as a sine qua non."

Hill was equally sharp about statistics. He argued that formal tests of significance "contribute nothing to the proof of our hypothesis", and serve only to remind us of the effects that the play of chance can create. That is a striking position from the man who did more than almost anyone to bring statistical method into medicine, and it is the correct antidote to a research culture that treats a p-value as a causal verdict.

The nine, translated for product research

Hill viewpointThe question it asksWhat it looks like in product research
StrengthHow large is the association?A 3x difference in churn demands less explaining away than a 4 percent difference. Weak associations are not disqualified, they just need more support elsewhere.
ConsistencyHas it been observed repeatedly, by different people, in different settings?The same pattern appears in the survey, in the interviews, in the support tickets, in two regions, and in another team's analysis.
SpecificityIs the effect specific to this exposure and this outcome?The onboarding problem predicts churn but not expansion or NPS. Hill considered this the weakest of the nine and said so.
TemporalityDid the cause come before the effect?The only one Hill called indispensable. Also the one destroyed by immortal time bias.
Biological gradientIs there a dose-response relationship?More exposure, more effect: accounts with three failed imports churn more than accounts with one. Very persuasive when present.
PlausibilityIs there a believable mechanism?Can you state, in a sentence, how the cause produces the effect in a human being's actual experience?
CoherenceDoes it fit what else is known?The explanation does not require your support volume, sales notes and usage data to all be wrong.
ExperimentDoes intervening change the outcome?A test, a staged rollout, a fix shipped to one segment. One of nine, not nine of nine.
AnalogyHas something similar been established elsewhere?The same failure caused the same outcome in an adjacent product or an earlier release.

Temporality carries a special status: an effect cannot precede its cause, so a failure here is fatal rather than merely weakening. That is exactly why the two sibling articles in this cluster matter so much. Immortal time bias manufactures a false temporal ordering out of misaligned clocks, and surveillance bias manufactures false strength and false consistency out of unequal looking. They are not separate topics from Hill. They are two specific ways of failing his viewpoints while appearing to satisfy them.

The anti-checklist rule

Here is the failure mode this framework invites, and it is worth naming before you use it: a team scores its finding, gets six of nine, and declares causation established. That is precisely what Hill said the viewpoints cannot do. None brings indisputable evidence; none is required.

Use them the other way round. Each viewpoint is not a box to tick but a specific alternative explanation you have or have not ruled out. Reverse each one into a disconfirming question:

ViewpointThe disconfirming question to ask instead
StrengthIf this were confounded rather than causal, how large would the association look?
ConsistencyWhich settings have I looked in where I would expect not to see it? Have I looked?
TemporalityWhat would I see if the arrow ran the other way, or if my clocks were misaligned?
Biological gradientIs there a dose-response, and if not, what mechanism explains an all-or-nothing effect?
PlausibilityCan I state the mechanism in one sentence that a customer would recognise?
CoherenceWhat existing evidence would have to be wrong for this to be true?
ExperimentWhat is the smallest intervention that would falsify this?
AnalogyWhere has a similar cause failed to produce this effect?

A finding that survives eight attempts to kill it is in a different epistemic position from one that scored eight ticks, even though the arithmetic looks the same. The disconfirming version also fixes the incentive problem: a checklist rewards a researcher for finding support, while a refutation list rewards them for finding the hole before a stakeholder does.

Three of the nine can only come from talking to people

This is the practical point that most quantitative teams miss. Look at the list again: plausibility, coherence and analogy are all judgements about mechanism and context. No dashboard emits them. They require someone to be able to say how and why the cause produces the effect in a person's actual experience.

Which means a research programme that can only run experiments and query event data is structurally incapable of addressing a third of Hill's viewpoints. It can measure strength, check temporality if the instrumentation is right, look for a gradient, and run the experiment where one is possible. It cannot supply a mechanism. Mechanism comes from asking people what happened and why, and then testing whether the account they give holds up across many people.

That is the honest case for qualitative research inside a causal argument. Not that interviews prove causation, which they do not, but that they supply three of the nine viewpoints that nothing else can supply, and they are the fastest route to falsifying a wrong mechanism.

The modern approach: how Koji helps

The reason teams skip the observational-causal discipline is cost. Building a Hill-style case traditionally means triangulating across several studies, several methods and several segments, which is weeks of recruiting and moderating for a question that a stakeholder wants answered on Thursday. So the team ships the correlation instead.

Koji changes the arithmetic on the viewpoints that need people:

  • Consistency becomes affordable. Consistency means the same result in different populations, and that is a sample-size and speed problem. Running the same brief across three segments and two regions simultaneously with AI-moderated interviews takes days rather than a quarter. Manual research forces you to pick one segment and hope.
  • Plausibility gets tested, not assumed. Instead of a researcher proposing a mechanism and a team nodding, you can put the mechanism to 60 customers and see whether they recognise it. A mechanism that only the product team believes is the most common cause of a wrong roadmap.
  • Temporality is established from the participant's account. Structured questions do the work here. Use single_choice to fix the ordering of two events, scale to size the contribution of each candidate cause, ranking to force a comparison across competing explanations, yes_no for clean gates, multiple_choice for the set of contributing factors, and open_ended with AI follow-up for the narrative that reveals a sequence the event log never captured. The structured questions guide covers combining them.
  • Biological gradient becomes measurable in qual. Recruit by exposure level, three failed imports versus one, and compare. This is a dose-response design and it is entirely feasible when fielding is cheap.
  • Automatic thematic analysis provides coherence evidence at scale. Coherence means the mechanism fits the rest of what you know, and thematic analysis across hundreds of transcripts is how you check whether the story holds outside the five calls you remember.
  • Customisable AI consultants let you encode the specific disconfirming questions above into the interview brief, so the study is designed to find the hole rather than to confirm the hypothesis.
  • Real-time reporting means you can stop when the mechanism has been falsified rather than completing a study you already know the answer to.

Against legacy survey tools the difference is structural. A SurveyMonkey field can measure strength and gradient but cannot probe a mechanism, so it cannot deliver plausibility or coherence. A traditional interview programme can deliver mechanism but not at the sample sizes and speeds that consistency demands. AI-moderated research is the first approach that delivers both halves within a single decision cycle, which is what makes a Hill-style case practical rather than aspirational. You do not need formal training in epidemiology to run this; you need a framework and a fast instrument.

A worked example

Claim: the failed-import experience causes churn.

ViewpointEvidence gatheredVerdict
StrengthAccounts with a failed import churn at 2.9x the base rateStrong, but confounding is plausible
TemporalityLandmark analysis at day 30 confirms failures precede cancellation; original chart had immortal time biasHolds after correction
GradientOne failure 1.4x, two 2.2x, three or more 3.8xClean dose-response
Plausibility42 of 60 interviewees describe abandoning setup and concluding the product could not handle their dataMechanism stated and recognised
ConsistencyPresent in SMB and mid-market, absent in enterprise where an onboarding engineer intervenesConsistent, with an explained exception
CoherenceSupport ticket themes and sales loss reasons alignNo contradiction
ExperimentImport diagnostics shipped to one segment; churn falls in that segment onlySupported
SpecificityPredicts churn but not NPS or expansionPresent, and Hill would remind you it is the weakest viewpoint
AnalogySame pattern followed the data-migration failures in a prior releaseSupported

That case is not proof. Hill would insist on the point. But it is a great deal more than a correlation, every viewpoint has been turned into a refutation attempt, and the enterprise exception actively strengthens it: the mechanism predicts where the effect should not appear, and it does not appear there.

Frequently asked questions

What are the Bradford Hill criteria?

They are nine viewpoints proposed by Sir Austin Bradford Hill in 1965 for judging whether an observed association is causal: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. Hill called them viewpoints rather than criteria, and stated that none can bring indisputable evidence and none can be required as a sine qua non.

Why use Bradford Hill instead of just running an A/B test?

Run the A/B test whenever you can; randomisation neutralises unknown confounders and nothing else does. Hill's framework is for the majority of product questions where randomisation is impossible or unethical, such as churn, outages, company size, or competitor actions. Hill developed it precisely because smoking could never be randomised.

Is the Bradford Hill framework a checklist?

No, and using it as one is the main way people misuse it. Hill explicitly said none of the nine is required and none is decisive. The productive use is to turn each viewpoint into a disconfirming question, treating each as a specific alternative explanation you have either ruled out or not.

Which of the nine viewpoints matters most?

Temporality, because an effect cannot precede its cause, so failing it is fatal rather than merely weakening. Biological gradient and strength carry a lot of weight when present. Hill regarded specificity as the weakest, since one cause can produce many effects. Note that temporality is exactly what immortal time bias destroys.

Can qualitative research contribute to causal claims?

Yes, and it is the only source for three of the nine. Plausibility, coherence and analogy are judgements about mechanism and context that no dashboard produces. Interviews do not prove causation on their own, but a causal case with no mechanism is weak, and mechanism only comes from asking people what happened and why.

Did Bradford Hill think statistical significance established causation?

Emphatically not. He wrote that tests of significance contribute nothing to the proof of a hypothesis, and serve only to remind us of the effects the play of chance can create. He regarded the causal judgement as something formed from the whole weight of evidence, not delivered by a single statistical test.

Related Resources

Related Articles

Case-Control Research: How to Study Churn and Lost Deals Without Fooling Yourself (2026)

Every churn interview and win-loss study is a case-control design, whether or not anyone says so. Epidemiology has spent seventy-five years learning how these studies go wrong - control selection, admission bias, recall bias, and base rates.

Correlation vs. Causation: Why Your Metrics Lie (and How to Find the Real Why)

A practical guide to correlation versus causation for product and research teams: why the two get confused, the classic traps, how to establish real causation, and how qualitative interviews reveal the mechanism behind the numbers.

Evidence Synthesis: How to Combine Findings Across Multiple Research Studies (2026)

Most teams have dozens of studies and no way to say what they collectively know. Evidence synthesis is the discipline of pooling findings across studies into a single rated conclusion - adapted from GRADE and systematic review practice for product research.

Immortal Time Bias: Why Feature Adopters Always Look More Loyal Than They Are (2026)

Immortal time bias makes every feature-adoption retention chart overstate the feature. Learn how the bias works, why product data is the worst case, and the three fixes.

How to Write a Research Hypothesis: A Step-by-Step Guide for Product & UX Teams

Master the art of writing testable research hypotheses. Learn the if-then-because format, null vs alternative hypotheses, common pitfalls, and how AI-native research turns hypotheses into validated learnings in days, not months.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Surveillance Bias: Why the Team That Measures Best Looks Worst (2026)

The harder you look, the more you find. Surveillance and lead time bias make well-instrumented teams look worse and useless interventions look effective. Here is how to tell the difference.

The Complete Guide to Thematic Analysis

Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.