Back to docs
Research Methods

Analysis of Competing Hypotheses: How to Test What Your Research Actually Supports

Most evidence that supports your favorite explanation also supports the ones you never wrote down. ACH is the matrix method that finds the evidence which actually discriminates.

Analysis of Competing Hypotheses (ACH) is a structured method for evaluating several explanations at once instead of building a case for the one you thought of first. You list every plausible hypothesis, list your evidence, and score each piece of evidence against each hypothesis in a matrix - not by asking "does this support my theory?" but by asking "if this hypothesis were true, how likely is it I would be seeing this?" The hypothesis you should believe is usually not the one with the most evidence for it. It is the one with the least evidence against it.

The short answer

The reason research teams reach confident wrong conclusions is rarely a shortage of evidence. It is that most of the evidence they collected is consistent with several explanations at once, and nobody checked. Ten customers say onboarding is confusing. That is consistent with "the onboarding is confusing." It is equally consistent with "we recruited the wrong customers," "the product is aimed at a job these people do not have," and "onboarding is fine but the pricing page over-promised." Evidence that fits every hypothesis has told you nothing, and it feels exactly as persuasive as evidence that fits only one.

ACH gives you a mechanical way to separate the two. It comes out of intelligence analysis - Richards J. Heuer, Jr. developed it at the CIA and published it in Psychology of Intelligence Analysis (1999), where it appears as the flagship structured analytic technique. It transfers to product research almost unchanged, because the underlying problem is identical: a smart person with partial information, a deadline, and a favored explanation.

The failure mode ACH is built against

The default human procedure is what Herbert Simon called satisficing: form a working hypothesis early, then evaluate incoming information against it. It is fast, and it is right most of the time, which is exactly why it survives. Heuer's argument is that it fails in a specific and predictable way - it never establishes whether the supporting evidence discriminates. You accumulate a pile of observations consistent with your theory and mistake the size of the pile for the strength of the case.

Two consequences follow, and both are visible in ordinary research readouts.

  • Confirmatory evidence is cheap and therefore uninformative. If a finding would look the same under three explanations, it cannot move you between them, no matter how many participants said it.
  • The absence of evidence goes unnoticed. Satisficing gives you no prompt to ask what you should be seeing if your hypothesis were true. A hypothesis that predicts something you cannot find in your data is in trouble, and you will never notice unless you wrote the prediction down first.

Diagnosticity: the one concept worth taking

Diagnosticity is the property that makes a piece of evidence useful. A finding is diagnostic to the degree that it is consistent with some hypotheses and inconsistent with others. A finding that is consistent with all of them has zero diagnostic value and belongs out of the analysis entirely, however vivid the quote.

This inverts how research evidence is usually assembled. A typical readout is sorted by how strongly it supports the conclusion. An ACH matrix is sorted by how much it separates the conclusions. Those two orderings can be nearly opposite, and the second one is the one that carries information.

The eight steps, adapted for product research

StepHeuer's versionWhat it means in a product study
1Identify the possible hypothesesWrite down every plausible explanation, including the boring ones (instrumentation, sampling, seasonality)
2List significant evidence for and against eachInclude what you observed and what you expected to observe but did not
3Build the matrix and assess diagnosticityHypotheses across the top, evidence down the side; mark consistent / inconsistent for each cell
4Refine: delete non-diagnostic evidenceAny row that reads "consistent" all the way across gets struck out
5Draw tentative conclusions by disprovingRank hypotheses by weight of inconsistent evidence, lowest first
6Test sensitivity to critical itemsAsk which one or two findings are carrying the conclusion, and how confident you are in each
7Report the relative likelihood of all hypothesesPresent the runners-up and why they lost, not only the winner
8Identify milestones for future observationWrite down what you would need to see to change your mind

Step 4 is where most of the value is realized and where most teams flinch, because the rows you strike out are usually the most quotable ones. Step 8 is the cheapest and most often skipped: a single sentence naming the observation that would falsify your conclusion converts a static readout into something testable next quarter.

A worked example

A B2B team sees weekly active usage fall 18 percent in a segment. The instinct is a single hypothesis: the new navigation broke something. ACH forces four.

  • H1: The redesigned navigation made a core workflow harder to reach.
  • H2: A large customer churned or reduced seats, and the segment average is composition, not behavior.
  • H3: Seasonal - the segment is education-heavy and this is a term break.
  • H4: Instrumentation changed and we are now counting differently.
EvidenceH1 navH2 compositionH3 seasonalH4 instrumentation
Usage fell 18 percentConsistentConsistentConsistentConsistent
Interviewees complain navigation is confusingConsistentConsistentConsistentConsistent
Decline started the exact day of the releaseConsistentInconsistentInconsistentConsistent
Per-account median usage is flatInconsistentConsistentInconsistentInconsistent
Same decline appears in last year's dataInconsistentInconsistentConsistentInconsistent

The first two rows are the ones a normal readout would lead with, and they are worthless - they are consistent with everything. The third row kills two hypotheses. The fourth row, a single boring line of analysis, is inconsistent with the hypothesis everyone arrived believing. Note that the navigation complaints are perfectly real and still do not discriminate: people find navigation confusing in every study ever run, including the ones where usage went up.

Under step 5 you rank by evidence against: H2 has one inconsistency, H1 has two, H3 has two, H4 has two. H2 leads, and the next study is a composition analysis, not a usability test.

The honest part: does ACH actually debias you?

Most articles about ACH stop at the matrix. The empirical record deserves to be stated plainly, because it changes how you should use the technique.

ACH was designed to counter confirmation bias, and the studies that have tested that claim have not supported it. Dhami, Belton and Mandel randomly assigned 50 intelligence analysts to use ACH or not on a hypothesis-testing task with probabilistic ground truth, and reported in Applied Cognitive Psychology (2019) that ACH-trained analysts did not follow all of the steps, that evidence for reducing confirmation bias was mixed, and that ACH may increase judgement inconsistency and error. Whitesmith's experimental work, summarized in Cognitive Bias in Intelligence Analysis (Edinburgh University Press, 2020), found no mitigation of confirmation bias or serial-position effects. Related work by Mandel and colleagues and by Karvetski and colleagues similarly found no advantage over an unstructured control.

So the debiasing claim is not supported. What survives is narrower and still worth your time:

  • Hypothesis generation. Being made to write down four explanations produces explanations you did not have. This is the benefit analysts consistently report, and it does not depend on the matrix scoring being correct.
  • Evidence auditing. The strike-out-the-non-diagnostic-rows step is a genuinely useful filter that nothing else in the standard research toolkit performs.
  • Externalization. The matrix makes your reasoning inspectable by someone else, which is a precondition for anyone disagreeing with you productively.

Treat ACH as a thinking aid and a communication artifact, not as a bias vaccine. If your goal is specifically to stop your expectations from steering your conclusions, the intervention with better evidence behind it is blind analysis - remove the information that lets you know which answer you are getting - combined with a pre-committed analysis plan.

Common mistakes

MistakeWhy it failsFix
Hypotheses that are not mutually exclusiveOverlapping hypotheses cannot be separated by any evidenceRewrite so at most one can be the primary driver
Only "real" hypotheses, no boring onesInstrumentation, sampling and seasonality are the most common true answersAlways include at least one measurement-artifact hypothesis
Scoring "supports" instead of "consistent with"Reintroduces the confirmation logic ACH is meant to replaceAsk: if this hypothesis were true, would I expect to see this?
Keeping the vivid non-diagnostic quoteIt feels like evidence and moves nothingStrike any row that is consistent across all columns
Reporting only the winnerHides how close the race wasReport the runners-up and the evidence that eliminated them
Never writing step 8The conclusion becomes unfalsifiableOne sentence: what would change my mind

The modern approach: making the evidence side cheap

The bottleneck in ACH has never been the matrix. It is that populating the evidence column properly requires evidence you usually do not have. You can generate four hypotheses in ten minutes and then discover that three of them are untestable with the data on hand, at which point the honest matrix is mostly blank and the team defaults back to the hypothesis they can argue for.

This is where the economics of research actually decide the analysis. Koji is built to close that gap:

  • Targeted follow-up studies in days, not weeks. Step 8 - "what would change my mind" - is only useful if you can go and look. AI-moderated interviews mean a hypothesis-discriminating study is a two-day exercise rather than a next-quarter commitment.
  • Structured questions produce discriminating evidence by design. All six types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - can be written to separate hypotheses rather than confirm one. A ranking question that forces a choice between four candidate causes is diagnostic; an open_ended "what frustrates you?" almost never is.
  • Automatic thematic analysis gives you prevalence, not anecdote. Diagnosticity depends on knowing how many people said something and who did not, which is exactly what manual quote-pulling loses. See thematic analysis and how to analyze qualitative data.
  • Customizable AI consultants can be pointed at the rival explanation. Running a short study designed to break your leading hypothesis is the ACH step nobody does, because it is a study you have to justify. When it costs a day, it stops needing justification.

Legacy tooling pushes the other way. A survey platform optimized for a single dashboard encourages one hypothesis with a supporting chart. The matrix wants several hypotheses and the evidence that separates them, and that requires being able to go back and ask again.

Frequently asked questions

How many hypotheses should I list?

Three to six. Below three you are not really comparing; above six the matrix becomes unreadable and the hypotheses start overlapping. If you have more, group them - and always reserve one slot for a measurement-artifact explanation, which is the single most commonly omitted and most commonly correct hypothesis in product analytics.

What is the difference between ACH and just listing alternative explanations?

The matrix and the diagnosticity filter. Listing alternatives is a brainstorm; ACH forces you to score every piece of evidence against every hypothesis and then delete the evidence that does not discriminate. The deletion step is the part that changes conclusions, and it is the part a brainstorm never reaches.

Why rank by evidence against rather than evidence for?

Because confirming evidence is abundant and usually non-diagnostic, while disconfirming evidence is rare and decisive. A single solid inconsistency can eliminate a hypothesis that a dozen consistent observations appeared to support. This is the practical form of falsification, and it is why the most probable hypothesis tends to be the one with the least evidence against it.

Does ACH work for qualitative research?

Yes, and arguably better than for quantitative work, because qualitative evidence is where non-diagnostic material accumulates fastest. Use coded themes as your evidence rows and theme prevalence as the consistency judgment. Pair it with inter-rater reliability so the coding underneath the matrix is stable.

Is ACH proven to reduce bias?

No. Controlled studies including Dhami, Belton and Mandel (2019) and Whitesmith (2020) found mixed-to-no evidence that ACH reduces confirmation bias, and some evidence it increases judgement inconsistency. Use it for hypothesis generation, evidence auditing and making your reasoning inspectable. For debiasing specifically, blind analysis and a pre-committed plan have better support.

How long does an ACH matrix take?

Forty-five minutes to two hours for a real question, most of it in step 2 assembling the evidence honestly. If it takes ten minutes you almost certainly listed hypotheses that are not mutually exclusive and skipped the diagnosticity filter.

The bottom line

The value of ACH is not that it makes you unbiased - the evidence says it does not. It is that it makes the shape of your case visible: how many explanations you actually considered, which evidence separated them, and what would change your mind. Most research readouts cannot answer any of those three questions, and a one-page matrix answers all of them.

Related Resources

Related Articles

Blind Analysis: How to Analyze Research Before You Know the Answer

Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.

Confirmation Bias in User Research: How to Recognize and Eliminate It

Confirmation bias quietly corrupts user research by leading teams to hear what they already believe. Learn how it shows up in interviews and analysis, and the practical tactics — and AI moderation — that neutralize it.

Conflicting Research Findings: What to Do When Qualitative and Quantitative Data Disagree (2026)

When your interviews say one thing and your analytics say another, averaging them is the worst possible move. A step-by-step protocol for diagnosing and resolving conflicting research findings.

How to Analyze Qualitative Data: From Raw Interviews to Actionable Insights

A step-by-step guide to qualitative data analysis — from reviewing raw transcripts to synthesizing themes, generating insights, and presenting findings that teams act on.

P-Hacking and Researcher Degrees of Freedom: How Analytic Flexibility Manufactures Findings (2026)

Four ordinary analytic choices raise the false-positive rate from 5 percent to 61 percent. Learn what researcher degrees of freedom are, why the garden of forking paths catches honest researchers, and how a one-page pre-committed analysis plan fixes it without banning exploration.

Research Synthesis: How to Combine Multiple Studies Into Clear Insights

A practical guide to synthesizing findings across multiple research studies — using thematic synthesis, triangulation, and structured data aggregation to build compounding organizational knowledge.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Triangulation in Research: Combining Methods for Stronger, More Credible Insights (2026)

Triangulation is the practice of using multiple data sources, methods, researchers, or theories to validate a finding. Learn Denzin's four types, when to use each, and how AI-native research platforms make multi-method studies practical instead of aspirational.