Analysis of Competing Hypotheses: How to Test What Your Research Actually Supports
Most evidence that supports your favorite explanation also supports the ones you never wrote down. ACH is the matrix method that finds the evidence which actually discriminates.
Analysis of Competing Hypotheses (ACH) is a structured method for evaluating several explanations at once instead of building a case for the one you thought of first. You list every plausible hypothesis, list your evidence, and score each piece of evidence against each hypothesis in a matrix - not by asking "does this support my theory?" but by asking "if this hypothesis were true, how likely is it I would be seeing this?" The hypothesis you should believe is usually not the one with the most evidence for it. It is the one with the least evidence against it.
The short answer
The reason research teams reach confident wrong conclusions is rarely a shortage of evidence. It is that most of the evidence they collected is consistent with several explanations at once, and nobody checked. Ten customers say onboarding is confusing. That is consistent with "the onboarding is confusing." It is equally consistent with "we recruited the wrong customers," "the product is aimed at a job these people do not have," and "onboarding is fine but the pricing page over-promised." Evidence that fits every hypothesis has told you nothing, and it feels exactly as persuasive as evidence that fits only one.
ACH gives you a mechanical way to separate the two. It comes out of intelligence analysis - Richards J. Heuer, Jr. developed it at the CIA and published it in Psychology of Intelligence Analysis (1999), where it appears as the flagship structured analytic technique. It transfers to product research almost unchanged, because the underlying problem is identical: a smart person with partial information, a deadline, and a favored explanation.
The failure mode ACH is built against
The default human procedure is what Herbert Simon called satisficing: form a working hypothesis early, then evaluate incoming information against it. It is fast, and it is right most of the time, which is exactly why it survives. Heuer's argument is that it fails in a specific and predictable way - it never establishes whether the supporting evidence discriminates. You accumulate a pile of observations consistent with your theory and mistake the size of the pile for the strength of the case.
Two consequences follow, and both are visible in ordinary research readouts.
- Confirmatory evidence is cheap and therefore uninformative. If a finding would look the same under three explanations, it cannot move you between them, no matter how many participants said it.
- The absence of evidence goes unnoticed. Satisficing gives you no prompt to ask what you should be seeing if your hypothesis were true. A hypothesis that predicts something you cannot find in your data is in trouble, and you will never notice unless you wrote the prediction down first.
Diagnosticity: the one concept worth taking
Diagnosticity is the property that makes a piece of evidence useful. A finding is diagnostic to the degree that it is consistent with some hypotheses and inconsistent with others. A finding that is consistent with all of them has zero diagnostic value and belongs out of the analysis entirely, however vivid the quote.
This inverts how research evidence is usually assembled. A typical readout is sorted by how strongly it supports the conclusion. An ACH matrix is sorted by how much it separates the conclusions. Those two orderings can be nearly opposite, and the second one is the one that carries information.
The eight steps, adapted for product research
| Step | Heuer's version | What it means in a product study |
|---|---|---|
| 1 | Identify the possible hypotheses | Write down every plausible explanation, including the boring ones (instrumentation, sampling, seasonality) |
| 2 | List significant evidence for and against each | Include what you observed and what you expected to observe but did not |
| 3 | Build the matrix and assess diagnosticity | Hypotheses across the top, evidence down the side; mark consistent / inconsistent for each cell |
| 4 | Refine: delete non-diagnostic evidence | Any row that reads "consistent" all the way across gets struck out |
| 5 | Draw tentative conclusions by disproving | Rank hypotheses by weight of inconsistent evidence, lowest first |
| 6 | Test sensitivity to critical items | Ask which one or two findings are carrying the conclusion, and how confident you are in each |
| 7 | Report the relative likelihood of all hypotheses | Present the runners-up and why they lost, not only the winner |
| 8 | Identify milestones for future observation | Write down what you would need to see to change your mind |
Step 4 is where most of the value is realized and where most teams flinch, because the rows you strike out are usually the most quotable ones. Step 8 is the cheapest and most often skipped: a single sentence naming the observation that would falsify your conclusion converts a static readout into something testable next quarter.
A worked example
A B2B team sees weekly active usage fall 18 percent in a segment. The instinct is a single hypothesis: the new navigation broke something. ACH forces four.
- H1: The redesigned navigation made a core workflow harder to reach.
- H2: A large customer churned or reduced seats, and the segment average is composition, not behavior.
- H3: Seasonal - the segment is education-heavy and this is a term break.
- H4: Instrumentation changed and we are now counting differently.
| Evidence | H1 nav | H2 composition | H3 seasonal | H4 instrumentation |
|---|---|---|---|---|
| Usage fell 18 percent | Consistent | Consistent | Consistent | Consistent |
| Interviewees complain navigation is confusing | Consistent | Consistent | Consistent | Consistent |
| Decline started the exact day of the release | Consistent | Inconsistent | Inconsistent | Consistent |
| Per-account median usage is flat | Inconsistent | Consistent | Inconsistent | Inconsistent |
| Same decline appears in last year's data | Inconsistent | Inconsistent | Consistent | Inconsistent |
The first two rows are the ones a normal readout would lead with, and they are worthless - they are consistent with everything. The third row kills two hypotheses. The fourth row, a single boring line of analysis, is inconsistent with the hypothesis everyone arrived believing. Note that the navigation complaints are perfectly real and still do not discriminate: people find navigation confusing in every study ever run, including the ones where usage went up.
Under step 5 you rank by evidence against: H2 has one inconsistency, H1 has two, H3 has two, H4 has two. H2 leads, and the next study is a composition analysis, not a usability test.
The honest part: does ACH actually debias you?
Most articles about ACH stop at the matrix. The empirical record deserves to be stated plainly, because it changes how you should use the technique.
ACH was designed to counter confirmation bias, and the studies that have tested that claim have not supported it. Dhami, Belton and Mandel randomly assigned 50 intelligence analysts to use ACH or not on a hypothesis-testing task with probabilistic ground truth, and reported in Applied Cognitive Psychology (2019) that ACH-trained analysts did not follow all of the steps, that evidence for reducing confirmation bias was mixed, and that ACH may increase judgement inconsistency and error. Whitesmith's experimental work, summarized in Cognitive Bias in Intelligence Analysis (Edinburgh University Press, 2020), found no mitigation of confirmation bias or serial-position effects. Related work by Mandel and colleagues and by Karvetski and colleagues similarly found no advantage over an unstructured control.
So the debiasing claim is not supported. What survives is narrower and still worth your time:
- Hypothesis generation. Being made to write down four explanations produces explanations you did not have. This is the benefit analysts consistently report, and it does not depend on the matrix scoring being correct.
- Evidence auditing. The strike-out-the-non-diagnostic-rows step is a genuinely useful filter that nothing else in the standard research toolkit performs.
- Externalization. The matrix makes your reasoning inspectable by someone else, which is a precondition for anyone disagreeing with you productively.
Treat ACH as a thinking aid and a communication artifact, not as a bias vaccine. If your goal is specifically to stop your expectations from steering your conclusions, the intervention with better evidence behind it is blind analysis - remove the information that lets you know which answer you are getting - combined with a pre-committed analysis plan.
Common mistakes
| Mistake | Why it fails | Fix |
|---|---|---|
| Hypotheses that are not mutually exclusive | Overlapping hypotheses cannot be separated by any evidence | Rewrite so at most one can be the primary driver |
| Only "real" hypotheses, no boring ones | Instrumentation, sampling and seasonality are the most common true answers | Always include at least one measurement-artifact hypothesis |
| Scoring "supports" instead of "consistent with" | Reintroduces the confirmation logic ACH is meant to replace | Ask: if this hypothesis were true, would I expect to see this? |
| Keeping the vivid non-diagnostic quote | It feels like evidence and moves nothing | Strike any row that is consistent across all columns |
| Reporting only the winner | Hides how close the race was | Report the runners-up and the evidence that eliminated them |
| Never writing step 8 | The conclusion becomes unfalsifiable | One sentence: what would change my mind |
The modern approach: making the evidence side cheap
The bottleneck in ACH has never been the matrix. It is that populating the evidence column properly requires evidence you usually do not have. You can generate four hypotheses in ten minutes and then discover that three of them are untestable with the data on hand, at which point the honest matrix is mostly blank and the team defaults back to the hypothesis they can argue for.
This is where the economics of research actually decide the analysis. Koji is built to close that gap:
- Targeted follow-up studies in days, not weeks. Step 8 - "what would change my mind" - is only useful if you can go and look. AI-moderated interviews mean a hypothesis-discriminating study is a two-day exercise rather than a next-quarter commitment.
- Structured questions produce discriminating evidence by design. All six types -
open_ended,scale,single_choice,multiple_choice,ranking, andyes_no- can be written to separate hypotheses rather than confirm one. Arankingquestion that forces a choice between four candidate causes is diagnostic; anopen_ended"what frustrates you?" almost never is. - Automatic thematic analysis gives you prevalence, not anecdote. Diagnosticity depends on knowing how many people said something and who did not, which is exactly what manual quote-pulling loses. See thematic analysis and how to analyze qualitative data.
- Customizable AI consultants can be pointed at the rival explanation. Running a short study designed to break your leading hypothesis is the ACH step nobody does, because it is a study you have to justify. When it costs a day, it stops needing justification.
Legacy tooling pushes the other way. A survey platform optimized for a single dashboard encourages one hypothesis with a supporting chart. The matrix wants several hypotheses and the evidence that separates them, and that requires being able to go back and ask again.
Frequently asked questions
How many hypotheses should I list?
Three to six. Below three you are not really comparing; above six the matrix becomes unreadable and the hypotheses start overlapping. If you have more, group them - and always reserve one slot for a measurement-artifact explanation, which is the single most commonly omitted and most commonly correct hypothesis in product analytics.
What is the difference between ACH and just listing alternative explanations?
The matrix and the diagnosticity filter. Listing alternatives is a brainstorm; ACH forces you to score every piece of evidence against every hypothesis and then delete the evidence that does not discriminate. The deletion step is the part that changes conclusions, and it is the part a brainstorm never reaches.
Why rank by evidence against rather than evidence for?
Because confirming evidence is abundant and usually non-diagnostic, while disconfirming evidence is rare and decisive. A single solid inconsistency can eliminate a hypothesis that a dozen consistent observations appeared to support. This is the practical form of falsification, and it is why the most probable hypothesis tends to be the one with the least evidence against it.
Does ACH work for qualitative research?
Yes, and arguably better than for quantitative work, because qualitative evidence is where non-diagnostic material accumulates fastest. Use coded themes as your evidence rows and theme prevalence as the consistency judgment. Pair it with inter-rater reliability so the coding underneath the matrix is stable.
Is ACH proven to reduce bias?
No. Controlled studies including Dhami, Belton and Mandel (2019) and Whitesmith (2020) found mixed-to-no evidence that ACH reduces confirmation bias, and some evidence it increases judgement inconsistency. Use it for hypothesis generation, evidence auditing and making your reasoning inspectable. For debiasing specifically, blind analysis and a pre-committed plan have better support.
How long does an ACH matrix take?
Forty-five minutes to two hours for a real question, most of it in step 2 assembling the evidence honestly. If it takes ten minutes you almost certainly listed hypotheses that are not mutually exclusive and skipped the diagnosticity filter.
The bottom line
The value of ACH is not that it makes you unbiased - the evidence says it does not. It is that it makes the shape of your case visible: how many explanations you actually considered, which evidence separated them, and what would change your mind. Most research readouts cannot answer any of those three questions, and a one-page matrix answers all of them.
Related Resources
- Blind Analysis - the intervention with better evidence for actually reducing analyst bias
- P-Hacking and Researcher Degrees of Freedom - what happens when the hypothesis is chosen after the data
- Confirmation Bias in User Research - the bias ACH was designed against
- Structured Questions Guide - writing questions that discriminate between hypotheses, across all six types
- Conflicting Research Findings - what to do when qualitative and quantitative evidence disagree
- Triangulation in Research - combining sources so more of your evidence is diagnostic
- Research Synthesis Guide - assembling findings without building a case
Related Articles
Blind Analysis: How to Analyze Research Before You Know the Answer
Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.
Confirmation Bias in User Research: How to Recognize and Eliminate It
Confirmation bias quietly corrupts user research by leading teams to hear what they already believe. Learn how it shows up in interviews and analysis, and the practical tactics — and AI moderation — that neutralize it.
Conflicting Research Findings: What to Do When Qualitative and Quantitative Data Disagree (2026)
When your interviews say one thing and your analytics say another, averaging them is the worst possible move. A step-by-step protocol for diagnosing and resolving conflicting research findings.
How to Analyze Qualitative Data: From Raw Interviews to Actionable Insights
A step-by-step guide to qualitative data analysis — from reviewing raw transcripts to synthesizing themes, generating insights, and presenting findings that teams act on.
P-Hacking and Researcher Degrees of Freedom: How Analytic Flexibility Manufactures Findings (2026)
Four ordinary analytic choices raise the false-positive rate from 5 percent to 61 percent. Learn what researcher degrees of freedom are, why the garden of forking paths catches honest researchers, and how a one-page pre-committed analysis plan fixes it without banning exploration.
Research Synthesis: How to Combine Multiple Studies Into Clear Insights
A practical guide to synthesizing findings across multiple research studies — using thematic synthesis, triangulation, and structured data aggregation to build compounding organizational knowledge.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Triangulation in Research: Combining Methods for Stronger, More Credible Insights (2026)
Triangulation is the practice of using multiple data sources, methods, researchers, or theories to validate a finding. Learn Denzin's four types, when to use each, and how AI-native research platforms make multi-method studies practical instead of aspirational.