{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-13T09:51:41.765Z"},"content":[{"type":"documentation","id":"4e8aa9dc-f5b8-4d05-8c3b-5cf4dcb51425","slug":"bradford-hill-criteria-product-research","title":"The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment (2026)","url":"https://www.koji.so/docs/bradford-hill-criteria-product-research","summary":"The Bradford Hill criteria are nine viewpoints proposed in 1965 for judging whether an association is causal: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. Hill called them viewpoints, not criteria, and wrote that none brings indisputable evidence and none is a sine qua non. He also held that significance tests contribute nothing to proof. The framework applies where randomisation is impossible, which covers most product questions. Only one of nine is experiment; plausibility, coherence and analogy require qualitative mechanism evidence. Best used by converting each viewpoint into a disconfirming question.","content":"# The Bradford Hill Criteria: Making Causal Claims When You Cannot Run the Experiment\n\n**Answer first:** The standard advice for establishing causation is to run a randomised experiment. It is correct advice, and it is unavailable for most of the questions product teams actually care about. You cannot randomly assign which customers churn, which companies are enterprises, which users hit the outage, or which buyer holds the budget. In 1965 Austin Bradford Hill set out nine viewpoints for weighing causation from observational evidence, precisely because the exposure in question, smoking, could never be randomised. Only one of the nine is experiment. This is the framework for the other cases, and it is deliberately not a checklist.\n\n## When the gold standard is not on the menu\n\nRandomised experimentation deserves its reputation. Where you can run an A/B test, run one: random assignment neutralises confounders you have not even thought of, which is something no amount of analysis can do. That case is covered in [correlation vs causation](/docs/correlation-vs-causation-research), and nothing here contradicts it.\n\nBut make an honest list of the causal questions your team asked last quarter:\n\n- Does poor onboarding cause churn?\n- Did the outage cost us the renewal?\n- Do teams that adopt our integration expand faster because of the integration?\n- Does pricing confusion cause abandoned trials?\n- Did the competitor's launch cause our win rate to fall?\n\nNot one of these can be randomised. You cannot assign customers to have a bad onboarding, or to experience an outage, or to be in a market where a competitor launched. Some are impossible, some are unethical, and some are merely commercially unthinkable. The experiment is off the table, the question still needs an answer, and the decision gets made either way.\n\nThe choice is not between experimental evidence and observational evidence. It is between disciplined observational evidence and somebody's confident hunch.\n\n## Hill in 1965, and what he actually said\n\nSir Austin Bradford Hill delivered his inaugural presidential address to the Section of Occupational Medicine of the Royal Society of Medicine in 1965, published as \"The Environment and Disease: Association or Causation?\" in the *Proceedings of the Royal Society of Medicine* (58(5):295-300). Hill had already co-authored, with Richard Doll, the case-control work that established the link between smoking and lung cancer, and he had done it without a single randomised trial because no such trial was possible.\n\nHe proposed nine viewpoints, in this order: **strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy.** In his own words: \"Here, then, are nine different viewpoints from all of which we should study association before we cry causation.\"\n\nTwo things about this are routinely got wrong.\n\nFirst, Hill never called them criteria. He called them viewpoints. The word \"criteria\" was applied by later writers and it changed the meaning: criteria sound like a standard you either meet or fail, whereas viewpoints are angles from which to look.\n\nSecond, and this is the sentence that most people citing Hill have never read: \"None of my nine viewpoints can bring indisputable evidence for or against the cause-and-effect hypothesis and none can be required as a *sine qua non*.\"\n\nHill was equally sharp about statistics. He argued that formal tests of significance \"contribute nothing to the proof of our hypothesis\", and serve only to remind us of the effects that the play of chance can create. That is a striking position from the man who did more than almost anyone to bring statistical method into medicine, and it is the correct antidote to a research culture that treats a p-value as a causal verdict.\n\n## The nine, translated for product research\n\n| Hill viewpoint | The question it asks | What it looks like in product research |\n| --- | --- | --- |\n| Strength | How large is the association? | A 3x difference in churn demands less explaining away than a 4 percent difference. Weak associations are not disqualified, they just need more support elsewhere. |\n| Consistency | Has it been observed repeatedly, by different people, in different settings? | The same pattern appears in the survey, in the interviews, in the support tickets, in two regions, and in another team's analysis. |\n| Specificity | Is the effect specific to this exposure and this outcome? | The onboarding problem predicts churn but not expansion or NPS. Hill considered this the weakest of the nine and said so. |\n| Temporality | Did the cause come before the effect? | The only one Hill called indispensable. Also the one destroyed by [immortal time bias](/docs/immortal-time-bias-retention-analysis). |\n| Biological gradient | Is there a dose-response relationship? | More exposure, more effect: accounts with three failed imports churn more than accounts with one. Very persuasive when present. |\n| Plausibility | Is there a believable mechanism? | Can you state, in a sentence, how the cause produces the effect in a human being's actual experience? |\n| Coherence | Does it fit what else is known? | The explanation does not require your support volume, sales notes and usage data to all be wrong. |\n| Experiment | Does intervening change the outcome? | A test, a staged rollout, a fix shipped to one segment. One of nine, not nine of nine. |\n| Analogy | Has something similar been established elsewhere? | The same failure caused the same outcome in an adjacent product or an earlier release. |\n\nTemporality carries a special status: an effect cannot precede its cause, so a failure here is fatal rather than merely weakening. That is exactly why the two sibling articles in this cluster matter so much. Immortal time bias manufactures a false temporal ordering out of misaligned clocks, and [surveillance bias](/docs/surveillance-bias-detection-research) manufactures false strength and false consistency out of unequal looking. They are not separate topics from Hill. They are two specific ways of failing his viewpoints while appearing to satisfy them.\n\n## The anti-checklist rule\n\nHere is the failure mode this framework invites, and it is worth naming before you use it: a team scores its finding, gets six of nine, and declares causation established. That is precisely what Hill said the viewpoints cannot do. None brings indisputable evidence; none is required.\n\nUse them the other way round. Each viewpoint is not a box to tick but a **specific alternative explanation you have or have not ruled out**. Reverse each one into a disconfirming question:\n\n| Viewpoint | The disconfirming question to ask instead |\n| --- | --- |\n| Strength | If this were confounded rather than causal, how large would the association look? |\n| Consistency | Which settings have I looked in where I would expect *not* to see it? Have I looked? |\n| Temporality | What would I see if the arrow ran the other way, or if my clocks were misaligned? |\n| Biological gradient | Is there a dose-response, and if not, what mechanism explains an all-or-nothing effect? |\n| Plausibility | Can I state the mechanism in one sentence that a customer would recognise? |\n| Coherence | What existing evidence would have to be wrong for this to be true? |\n| Experiment | What is the smallest intervention that would falsify this? |\n| Analogy | Where has a similar cause failed to produce this effect? |\n\nA finding that survives eight attempts to kill it is in a different epistemic position from one that scored eight ticks, even though the arithmetic looks the same. The disconfirming version also fixes the incentive problem: a checklist rewards a researcher for finding support, while a refutation list rewards them for finding the hole before a stakeholder does.\n\n## Three of the nine can only come from talking to people\n\nThis is the practical point that most quantitative teams miss. Look at the list again: **plausibility, coherence and analogy** are all judgements about mechanism and context. No dashboard emits them. They require someone to be able to say how and why the cause produces the effect in a person's actual experience.\n\nWhich means a research programme that can only run experiments and query event data is structurally incapable of addressing a third of Hill's viewpoints. It can measure strength, check temporality if the instrumentation is right, look for a gradient, and run the experiment where one is possible. It cannot supply a mechanism. Mechanism comes from asking people what happened and why, and then testing whether the account they give holds up across many people.\n\nThat is the honest case for qualitative research inside a causal argument. Not that interviews prove causation, which they do not, but that they supply three of the nine viewpoints that nothing else can supply, and they are the fastest route to falsifying a wrong mechanism.\n\n## The modern approach: how Koji helps\n\nThe reason teams skip the observational-causal discipline is cost. Building a Hill-style case traditionally means triangulating across several studies, several methods and several segments, which is weeks of recruiting and moderating for a question that a stakeholder wants answered on Thursday. So the team ships the correlation instead.\n\nKoji changes the arithmetic on the viewpoints that need people:\n\n- **Consistency becomes affordable.** Consistency means the same result in different populations, and that is a sample-size and speed problem. Running the same brief across three segments and two regions simultaneously with AI-moderated interviews takes days rather than a quarter. Manual research forces you to pick one segment and hope.\n- **Plausibility gets tested, not assumed.** Instead of a researcher proposing a mechanism and a team nodding, you can put the mechanism to 60 customers and see whether they recognise it. A mechanism that only the product team believes is the most common cause of a wrong roadmap.\n- **Temporality is established from the participant's account.** Structured questions do the work here. Use `single_choice` to fix the ordering of two events, `scale` to size the contribution of each candidate cause, `ranking` to force a comparison across competing explanations, `yes_no` for clean gates, `multiple_choice` for the set of contributing factors, and `open_ended` with AI follow-up for the narrative that reveals a sequence the event log never captured. The [structured questions guide](/docs/structured-questions-guide) covers combining them.\n- **Biological gradient becomes measurable in qual.** Recruit by exposure level, three failed imports versus one, and compare. This is a dose-response design and it is entirely feasible when fielding is cheap.\n- **Automatic thematic analysis provides coherence evidence at scale.** Coherence means the mechanism fits the rest of what you know, and thematic analysis across hundreds of transcripts is how you check whether the story holds outside the five calls you remember.\n- **Customisable AI consultants** let you encode the specific disconfirming questions above into the interview brief, so the study is designed to find the hole rather than to confirm the hypothesis.\n- **Real-time reporting** means you can stop when the mechanism has been falsified rather than completing a study you already know the answer to.\n\nAgainst legacy survey tools the difference is structural. A SurveyMonkey field can measure strength and gradient but cannot probe a mechanism, so it cannot deliver plausibility or coherence. A traditional interview programme can deliver mechanism but not at the sample sizes and speeds that consistency demands. AI-moderated research is the first approach that delivers both halves within a single decision cycle, which is what makes a Hill-style case practical rather than aspirational. You do not need formal training in epidemiology to run this; you need a framework and a fast instrument.\n\n## A worked example\n\n**Claim:** the failed-import experience causes churn.\n\n| Viewpoint | Evidence gathered | Verdict |\n| --- | --- | --- |\n| Strength | Accounts with a failed import churn at 2.9x the base rate | Strong, but confounding is plausible |\n| Temporality | Landmark analysis at day 30 confirms failures precede cancellation; original chart had immortal time bias | Holds after correction |\n| Gradient | One failure 1.4x, two 2.2x, three or more 3.8x | Clean dose-response |\n| Plausibility | 42 of 60 interviewees describe abandoning setup and concluding the product could not handle their data | Mechanism stated and recognised |\n| Consistency | Present in SMB and mid-market, absent in enterprise where an onboarding engineer intervenes | Consistent, with an explained exception |\n| Coherence | Support ticket themes and sales loss reasons align | No contradiction |\n| Experiment | Import diagnostics shipped to one segment; churn falls in that segment only | Supported |\n| Specificity | Predicts churn but not NPS or expansion | Present, and Hill would remind you it is the weakest viewpoint |\n| Analogy | Same pattern followed the data-migration failures in a prior release | Supported |\n\nThat case is not proof. Hill would insist on the point. But it is a great deal more than a correlation, every viewpoint has been turned into a refutation attempt, and the enterprise exception actively strengthens it: the mechanism predicts where the effect should *not* appear, and it does not appear there.\n\n## Frequently asked questions\n\n### What are the Bradford Hill criteria?\n\nThey are nine viewpoints proposed by Sir Austin Bradford Hill in 1965 for judging whether an observed association is causal: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. Hill called them viewpoints rather than criteria, and stated that none can bring indisputable evidence and none can be required as a sine qua non.\n\n### Why use Bradford Hill instead of just running an A/B test?\n\nRun the A/B test whenever you can; randomisation neutralises unknown confounders and nothing else does. Hill's framework is for the majority of product questions where randomisation is impossible or unethical, such as churn, outages, company size, or competitor actions. Hill developed it precisely because smoking could never be randomised.\n\n### Is the Bradford Hill framework a checklist?\n\nNo, and using it as one is the main way people misuse it. Hill explicitly said none of the nine is required and none is decisive. The productive use is to turn each viewpoint into a disconfirming question, treating each as a specific alternative explanation you have either ruled out or not.\n\n### Which of the nine viewpoints matters most?\n\nTemporality, because an effect cannot precede its cause, so failing it is fatal rather than merely weakening. Biological gradient and strength carry a lot of weight when present. Hill regarded specificity as the weakest, since one cause can produce many effects. Note that temporality is exactly what immortal time bias destroys.\n\n### Can qualitative research contribute to causal claims?\n\nYes, and it is the only source for three of the nine. Plausibility, coherence and analogy are judgements about mechanism and context that no dashboard produces. Interviews do not prove causation on their own, but a causal case with no mechanism is weak, and mechanism only comes from asking people what happened and why.\n\n### Did Bradford Hill think statistical significance established causation?\n\nEmphatically not. He wrote that tests of significance contribute nothing to the proof of a hypothesis, and serve only to remind us of the effects the play of chance can create. He regarded the causal judgement as something formed from the whole weight of evidence, not delivered by a single statistical test.\n\n## Related Resources\n\n- [Structured questions guide](/docs/structured-questions-guide) - the six question types and how to use them to establish sequence and dose-response\n- [Immortal time bias](/docs/immortal-time-bias-retention-analysis) - the specific way temporality gets faked in retention data\n- [Surveillance bias](/docs/surveillance-bias-detection-research) - how unequal looking fakes strength and consistency\n- [Correlation vs causation](/docs/correlation-vs-causation-research) - the foundation, and the case for experiments where they are available\n- [Case-control research for churn and lost deals](/docs/case-control-research-churn-lost-deals) - the observational design most product teams need\n- [Evidence synthesis for research findings](/docs/evidence-synthesis-research-findings) - combining studies into a single defensible conclusion\n- [Research hypothesis](/docs/research-hypothesis) - stating a claim precisely enough to be refuted\n- [Thematic analysis guide](/docs/thematic-analysis-guide) - extracting mechanism from qualitative data at scale","category":"Research Methods","lastModified":"2026-08-13T03:24:41.060734+00:00","metaTitle":"Bradford Hill Criteria for Product Research: Causation Without Experiments","metaDescription":"Most product questions cannot be randomised. Learn Bradford Hill nine viewpoints, why they are not a checklist, and how to build a defensible causal case without an A/B test.","keywords":["bradford hill criteria","causal inference product research","association or causation","observational research causation","hill viewpoints","temporality","dose-response","causal claims without experiments"],"aiSummary":"The Bradford Hill criteria are nine viewpoints proposed in 1965 for judging whether an association is causal: strength, consistency, specificity, temporality, biological gradient, plausibility, coherence, experiment, and analogy. Hill called them viewpoints, not criteria, and wrote that none brings indisputable evidence and none is a sine qua non. He also held that significance tests contribute nothing to proof. The framework applies where randomisation is impossible, which covers most product questions. Only one of nine is experiment; plausibility, coherence and analogy require qualitative mechanism evidence. Best used by converting each viewpoint into a disconfirming question.","aiPrerequisites":["Understanding of the difference between correlation and causation","Familiarity with basic research design"],"aiLearningOutcomes":["Apply Hill nine viewpoints to an observational product finding","Explain why the viewpoints are not a checklist","Convert each viewpoint into a disconfirming question","Identify which viewpoints require qualitative evidence","Build a defensible causal case where no experiment is possible"],"aiDifficulty":"intermediate","aiEstimatedTime":"13 min"}],"pagination":{"total":1,"returned":1,"offset":0}}