The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment (2026)
Most of what matters in product research cannot be randomised. Bradford Hill nine viewpoints are the framework for building a defensible causal case without an A/B test.
The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment
Answer first: The standard advice for establishing causation is to run a randomised experiment. It is correct advice, and it is unavailable for most of the questions product teams actually care about. You cannot randomly assign which customers churn, which companies are enterprises, which users hit the outage, or which buyer holds the budget. In 1965 Austin Bradford Hill set out nine viewpoints for weighing causation from observational evidence, precisely because the exposure in question, smoking, could never be randomised. Only one of the nine is experiment. This is the framework for the other cases, and it is deliberately not a checklist.
When the gold standard is not on the menu
Randomised experimentation deserves its reputation. Where you can run an A/B test, run one: random assignment neutralises confounders you have not even thought of, which is something no amount of analysis can do. That case is covered in correlation vs causation, and nothing here contradicts it.
But make an honest list of the causal questions your team asked last quarter:
- Does poor onboarding cause churn?
- Did the outage cost us the renewal?
- Do teams that adopt our integration expand faster because of the integration?
- Does pricing confusion cause abandoned trials?
- Did the competitor's launch cause our win rate to fall?
Not one of these can be randomised. You cannot assign customers to have a bad onboarding, or to experience an outage, or to be in a market where a competitor launched. Some are impossible, some are unethical, and some are merely commercially unthinkable. The experiment is off the table, the question still needs an answer, and the decision gets made either way.
The choice is not between experimental evidence and observational evidence. It is between disciplined observational evidence and somebody's confident hunch.
Hill in 1965, and what he actually said
Sir Austin Bradford Hill delivered his inaugural presidential address to the Section of Occupational Medicine of the Royal Society of Medicine in 1965, published as "The Environment and Disease: Association or Causation?" in the Proceedings of the Royal Society of Medicine (58(5):295-300). Hill had already co-authored, with Richard Doll, the case-control work that established the link between smoking and lung cancer, and he had done it without a single randomised trial because no such trial was possible.
He proposed nine viewpoints, in this order: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. In his own words: "Here, then, are nine different viewpoints from all of which we should study association before we cry causation."
Two things about this are routinely got wrong.
First, Hill never called them criteria. He called them viewpoints. The word "criteria" was applied by later writers and it changed the meaning: criteria sound like a standard you either meet or fail, whereas viewpoints are angles from which to look.
Second, and this is the sentence that most people citing Hill have never read: "None of my nine viewpoints can bring indisputable evidence for or against the cause-and-effect hypothesis and none can be required as a sine qua non."
Hill was equally sharp about statistics. He argued that formal tests of significance "contribute nothing to the proof of our hypothesis", and serve only to remind us of the effects that the play of chance can create. That is a striking position from the man who did more than almost anyone to bring statistical method into medicine, and it is the correct antidote to a research culture that treats a p-value as a causal verdict.
The nine, translated for product research
| Hill viewpoint | The question it asks | What it looks like in product research |
|---|---|---|
| Strength | How large is the association? | A 3x difference in churn demands less explaining away than a 4 percent difference. Weak associations are not disqualified, they just need more support elsewhere. |
| Consistency | Has it been observed repeatedly, by different people, in different settings? | The same pattern appears in the survey, in the interviews, in the support tickets, in two regions, and in another team's analysis. |
| Specificity | Is the effect specific to this exposure and this outcome? | The onboarding problem predicts churn but not expansion or NPS. Hill considered this the weakest of the nine and said so. |
| Temporality | Did the cause come before the effect? | The only one Hill called indispensable. Also the one destroyed by immortal time bias. |
| Biological gradient | Is there a dose-response relationship? | More exposure, more effect: accounts with three failed imports churn more than accounts with one. Very persuasive when present. |
| Plausibility | Is there a believable mechanism? | Can you state, in a sentence, how the cause produces the effect in a human being's actual experience? |
| Coherence | Does it fit what else is known? | The explanation does not require your support volume, sales notes and usage data to all be wrong. |
| Experiment | Does intervening change the outcome? | A test, a staged rollout, a fix shipped to one segment. One of nine, not nine of nine. |
| Analogy | Has something similar been established elsewhere? | The same failure caused the same outcome in an adjacent product or an earlier release. |
Temporality carries a special status: an effect cannot precede its cause, so a failure here is fatal rather than merely weakening. That is exactly why the two sibling articles in this cluster matter so much. Immortal time bias manufactures a false temporal ordering out of misaligned clocks, and surveillance bias manufactures false strength and false consistency out of unequal looking. They are not separate topics from Hill. They are two specific ways of failing his viewpoints while appearing to satisfy them.
The anti-checklist rule
Here is the failure mode this framework invites, and it is worth naming before you use it: a team scores its finding, gets six of nine, and declares causation established. That is precisely what Hill said the viewpoints cannot do. None brings indisputable evidence; none is required.
Use them the other way round. Each viewpoint is not a box to tick but a specific alternative explanation you have or have not ruled out. Reverse each one into a disconfirming question:
| Viewpoint | The disconfirming question to ask instead |
|---|---|
| Strength | If this were confounded rather than causal, how large would the association look? |
| Consistency | Which settings have I looked in where I would expect not to see it? Have I looked? |
| Temporality | What would I see if the arrow ran the other way, or if my clocks were misaligned? |
| Biological gradient | Is there a dose-response, and if not, what mechanism explains an all-or-nothing effect? |
| Plausibility | Can I state the mechanism in one sentence that a customer would recognise? |
| Coherence | What existing evidence would have to be wrong for this to be true? |
| Experiment | What is the smallest intervention that would falsify this? |
| Analogy | Where has a similar cause failed to produce this effect? |
A finding that survives eight attempts to kill it is in a different epistemic position from one that scored eight ticks, even though the arithmetic looks the same. The disconfirming version also fixes the incentive problem: a checklist rewards a researcher for finding support, while a refutation list rewards them for finding the hole before a stakeholder does.
Three of the nine can only come from talking to people
This is the practical point that most quantitative teams miss. Look at the list again: plausibility, coherence and analogy are all judgements about mechanism and context. No dashboard emits them. They require someone to be able to say how and why the cause produces the effect in a person's actual experience.
Which means a research programme that can only run experiments and query event data is structurally incapable of addressing a third of Hill's viewpoints. It can measure strength, check temporality if the instrumentation is right, look for a gradient, and run the experiment where one is possible. It cannot supply a mechanism. Mechanism comes from asking people what happened and why, and then testing whether the account they give holds up across many people.
That is the honest case for qualitative research inside a causal argument. Not that interviews prove causation, which they do not, but that they supply three of the nine viewpoints that nothing else can supply, and they are the fastest route to falsifying a wrong mechanism.
The modern approach: how Koji helps
The reason teams skip the observational-causal discipline is cost. Building a Hill-style case traditionally means triangulating across several studies, several methods and several segments, which is weeks of recruiting and moderating for a question that a stakeholder wants answered on Thursday. So the team ships the correlation instead.
Koji changes the arithmetic on the viewpoints that need people:
- Consistency becomes affordable. Consistency means the same result in different populations, and that is a sample-size and speed problem. Running the same brief across three segments and two regions simultaneously with AI-moderated interviews takes days rather than a quarter. Manual research forces you to pick one segment and hope.
- Plausibility gets tested, not assumed. Instead of a researcher proposing a mechanism and a team nodding, you can put the mechanism to 60 customers and see whether they recognise it. A mechanism that only the product team believes is the most common cause of a wrong roadmap.
- Temporality is established from the participant's account. Structured questions do the work here. Use
single_choiceto fix the ordering of two events,scaleto size the contribution of each candidate cause,rankingto force a comparison across competing explanations,yes_nofor clean gates,multiple_choicefor the set of contributing factors, andopen_endedwith AI follow-up for the narrative that reveals a sequence the event log never captured. The structured questions guide covers combining them. - Biological gradient becomes measurable in qual. Recruit by exposure level, three failed imports versus one, and compare. This is a dose-response design and it is entirely feasible when fielding is cheap.
- Automatic thematic analysis provides coherence evidence at scale. Coherence means the mechanism fits the rest of what you know, and thematic analysis across hundreds of transcripts is how you check whether the story holds outside the five calls you remember.
- Customisable AI consultants let you encode the specific disconfirming questions above into the interview brief, so the study is designed to find the hole rather than to confirm the hypothesis.
- Real-time reporting means you can stop when the mechanism has been falsified rather than completing a study you already know the answer to.
Against legacy survey tools the difference is structural. A SurveyMonkey field can measure strength and gradient but cannot probe a mechanism, so it cannot deliver plausibility or coherence. A traditional interview programme can deliver mechanism but not at the sample sizes and speeds that consistency demands. AI-moderated research is the first approach that delivers both halves within a single decision cycle, which is what makes a Hill-style case practical rather than aspirational. You do not need formal training in epidemiology to run this; you need a framework and a fast instrument.
A worked example
Claim: the failed-import experience causes churn.
| Viewpoint | Evidence gathered | Verdict |
|---|---|---|
| Strength | Accounts with a failed import churn at 2.9x the base rate | Strong, but confounding is plausible |
| Temporality | Landmark analysis at day 30 confirms failures precede cancellation; original chart had immortal time bias | Holds after correction |
| Gradient | One failure 1.4x, two 2.2x, three or more 3.8x | Clean dose-response |
| Plausibility | 42 of 60 interviewees describe abandoning setup and concluding the product could not handle their data | Mechanism stated and recognised |
| Consistency | Present in SMB and mid-market, absent in enterprise where an onboarding engineer intervenes | Consistent, with an explained exception |
| Coherence | Support ticket themes and sales loss reasons align | No contradiction |
| Experiment | Import diagnostics shipped to one segment; churn falls in that segment only | Supported |
| Specificity | Predicts churn but not NPS or expansion | Present, and Hill would remind you it is the weakest viewpoint |
| Analogy | Same pattern followed the data-migration failures in a prior release | Supported |
That case is not proof. Hill would insist on the point. But it is a great deal more than a correlation, every viewpoint has been turned into a refutation attempt, and the enterprise exception actively strengthens it: the mechanism predicts where the effect should not appear, and it does not appear there.
Frequently asked questions
What are the Bradford Hill criteria?
They are nine viewpoints proposed by Sir Austin Bradford Hill in 1965 for judging whether an observed association is causal: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. Hill called them viewpoints rather than criteria, and stated that none can bring indisputable evidence and none can be required as a sine qua non.
Why use Bradford Hill instead of just running an A/B test?
Run the A/B test whenever you can; randomisation neutralises unknown confounders and nothing else does. Hill's framework is for the majority of product questions where randomisation is impossible or unethical, such as churn, outages, company size, or competitor actions. Hill developed it precisely because smoking could never be randomised.
Is the Bradford Hill framework a checklist?
No, and using it as one is the main way people misuse it. Hill explicitly said none of the nine is required and none is decisive. The productive use is to turn each viewpoint into a disconfirming question, treating each as a specific alternative explanation you have either ruled out or not.
Which of the nine viewpoints matters most?
Temporality, because an effect cannot precede its cause, so failing it is fatal rather than merely weakening. Biological gradient and strength carry a lot of weight when present. Hill regarded specificity as the weakest, since one cause can produce many effects. Note that temporality is exactly what immortal time bias destroys.
Can qualitative research contribute to causal claims?
Yes, and it is the only source for three of the nine. Plausibility, coherence and analogy are judgements about mechanism and context that no dashboard produces. Interviews do not prove causation on their own, but a causal case with no mechanism is weak, and mechanism only comes from asking people what happened and why.
Did Bradford Hill think statistical significance established causation?
Emphatically not. He wrote that tests of significance contribute nothing to the proof of a hypothesis, and serve only to remind us of the effects the play of chance can create. He regarded the causal judgement as something formed from the whole weight of evidence, not delivered by a single statistical test.
Related Resources
- Structured questions guide - the six question types and how to use them to establish sequence and dose-response
- Immortal time bias - the specific way temporality gets faked in retention data
- Surveillance bias - how unequal looking fakes strength and consistency
- Correlation vs causation - the foundation, and the case for experiments where they are available
- Case-control research for churn and lost deals - the observational design most product teams need
- Evidence synthesis for research findings - combining studies into a single defensible conclusion
- Research hypothesis - stating a claim precisely enough to be refuted
- Thematic analysis guide - extracting mechanism from qualitative data at scale
Related Articles
Case-Control Research: How to Study Churn and Lost Deals Without Fooling Yourself (2026)
Every churn interview and win-loss study is a case-control design, whether or not anyone says so. Epidemiology has spent seventy-five years learning how these studies go wrong - control selection, admission bias, recall bias, and base rates.
Correlation vs. Causation: Why Your Metrics Lie (and How to Find the Real Why)
A practical guide to correlation versus causation for product and research teams: why the two get confused, the classic traps, how to establish real causation, and how qualitative interviews reveal the mechanism behind the numbers.
Evidence Synthesis: How to Combine Findings Across Multiple Research Studies (2026)
Most teams have dozens of studies and no way to say what they collectively know. Evidence synthesis is the discipline of pooling findings across studies into a single rated conclusion - adapted from GRADE and systematic review practice for product research.
Immortal Time Bias: Why Feature Adopters Always Look More Loyal Than They Are (2026)
Immortal time bias makes every feature-adoption retention chart overstate the feature. Learn how the bias works, why product data is the worst case, and the three fixes.
How to Write a Research Hypothesis: A Step-by-Step Guide for Product & UX Teams
Master the art of writing testable research hypotheses. Learn the if-then-because format, null vs alternative hypotheses, common pitfalls, and how AI-native research turns hypotheses into validated learnings in days, not months.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Surveillance Bias: Why the Team That Measures Best Looks Worst (2026)
The harder you look, the more you find. Surveillance and lead time bias make well-instrumented teams look worse and useless interventions look effective. Here is how to tell the difference.
The Complete Guide to Thematic Analysis
Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.