{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-12T14:06:35.675Z"},"content":[{"type":"documentation","id":"be1096e0-27f6-4673-80d0-7e89ba5a5fe7","slug":"analysis-of-competing-hypotheses-research","title":"Analysis of Competing Hypotheses: How to Test What Your Research Actually Supports","url":"https://www.koji.so/docs/analysis-of-competing-hypotheses-research","summary":"Analysis of Competing Hypotheses (ACH), developed by Richards J. Heuer, Jr. at the CIA and published in Psychology of Intelligence Analysis (1999), evaluates several explanations simultaneously using a matrix of evidence against hypotheses. Its central concept is diagnosticity: evidence consistent with every hypothesis has zero value regardless of how persuasive it feels. Rank hypotheses by evidence against them, not for them. Controlled studies (Dhami, Belton and Mandel 2019; Whitesmith 2020) found mixed-to-no evidence that ACH reduces confirmation bias, so use it for hypothesis generation, evidence auditing and externalizing reasoning rather than as a debiasing method.","content":"Analysis of Competing Hypotheses (ACH) is a structured method for evaluating several explanations at once instead of building a case for the one you thought of first. You list every plausible hypothesis, list your evidence, and score each piece of evidence against each hypothesis in a matrix - not by asking \"does this support my theory?\" but by asking \"if this hypothesis were true, how likely is it I would be seeing this?\" The hypothesis you should believe is usually not the one with the most evidence for it. It is the one with the least evidence against it.\n\n## The short answer\n\nThe reason research teams reach confident wrong conclusions is rarely a shortage of evidence. It is that most of the evidence they collected is consistent with several explanations at once, and nobody checked. Ten customers say onboarding is confusing. That is consistent with \"the onboarding is confusing.\" It is equally consistent with \"we recruited the wrong customers,\" \"the product is aimed at a job these people do not have,\" and \"onboarding is fine but the pricing page over-promised.\" Evidence that fits every hypothesis has told you nothing, and it feels exactly as persuasive as evidence that fits only one.\n\nACH gives you a mechanical way to separate the two. It comes out of intelligence analysis - Richards J. Heuer, Jr. developed it at the CIA and published it in *Psychology of Intelligence Analysis* (1999), where it appears as the flagship structured analytic technique. It transfers to product research almost unchanged, because the underlying problem is identical: a smart person with partial information, a deadline, and a favored explanation.\n\n## The failure mode ACH is built against\n\nThe default human procedure is what Herbert Simon called satisficing: form a working hypothesis early, then evaluate incoming information against it. It is fast, and it is right most of the time, which is exactly why it survives. Heuer's argument is that it fails in a specific and predictable way - it never establishes whether the supporting evidence *discriminates*. You accumulate a pile of observations consistent with your theory and mistake the size of the pile for the strength of the case.\n\nTwo consequences follow, and both are visible in ordinary research readouts.\n\n- **Confirmatory evidence is cheap and therefore uninformative.** If a finding would look the same under three explanations, it cannot move you between them, no matter how many participants said it.\n- **The absence of evidence goes unnoticed.** Satisficing gives you no prompt to ask what you *should* be seeing if your hypothesis were true. A hypothesis that predicts something you cannot find in your data is in trouble, and you will never notice unless you wrote the prediction down first.\n\n## Diagnosticity: the one concept worth taking\n\nDiagnosticity is the property that makes a piece of evidence useful. A finding is diagnostic to the degree that it is consistent with some hypotheses and inconsistent with others. A finding that is consistent with all of them has zero diagnostic value and belongs out of the analysis entirely, however vivid the quote.\n\nThis inverts how research evidence is usually assembled. A typical readout is sorted by *how strongly it supports the conclusion*. An ACH matrix is sorted by *how much it separates the conclusions*. Those two orderings can be nearly opposite, and the second one is the one that carries information.\n\n## The eight steps, adapted for product research\n\n| Step | Heuer's version | What it means in a product study |\n| --- | --- | --- |\n| 1 | Identify the possible hypotheses | Write down every plausible explanation, including the boring ones (instrumentation, sampling, seasonality) |\n| 2 | List significant evidence for and against each | Include what you observed *and* what you expected to observe but did not |\n| 3 | Build the matrix and assess diagnosticity | Hypotheses across the top, evidence down the side; mark consistent / inconsistent for each cell |\n| 4 | Refine: delete non-diagnostic evidence | Any row that reads \"consistent\" all the way across gets struck out |\n| 5 | Draw tentative conclusions by disproving | Rank hypotheses by weight of *inconsistent* evidence, lowest first |\n| 6 | Test sensitivity to critical items | Ask which one or two findings are carrying the conclusion, and how confident you are in each |\n| 7 | Report the relative likelihood of all hypotheses | Present the runners-up and why they lost, not only the winner |\n| 8 | Identify milestones for future observation | Write down what you would need to see to change your mind |\n\nStep 4 is where most of the value is realized and where most teams flinch, because the rows you strike out are usually the most quotable ones. Step 8 is the cheapest and most often skipped: a single sentence naming the observation that would falsify your conclusion converts a static readout into something testable next quarter.\n\n## A worked example\n\nA B2B team sees weekly active usage fall 18 percent in a segment. The instinct is a single hypothesis: the new navigation broke something. ACH forces four.\n\n- **H1:** The redesigned navigation made a core workflow harder to reach.\n- **H2:** A large customer churned or reduced seats, and the segment average is composition, not behavior.\n- **H3:** Seasonal - the segment is education-heavy and this is a term break.\n- **H4:** Instrumentation changed and we are now counting differently.\n\n| Evidence | H1 nav | H2 composition | H3 seasonal | H4 instrumentation |\n| --- | --- | --- | --- | --- |\n| Usage fell 18 percent | Consistent | Consistent | Consistent | Consistent |\n| Interviewees complain navigation is confusing | Consistent | Consistent | Consistent | Consistent |\n| Decline started the exact day of the release | Consistent | Inconsistent | Inconsistent | Consistent |\n| Per-account median usage is flat | Inconsistent | Consistent | Inconsistent | Inconsistent |\n| Same decline appears in last year's data | Inconsistent | Inconsistent | Consistent | Inconsistent |\n\nThe first two rows are the ones a normal readout would lead with, and they are worthless - they are consistent with everything. The third row kills two hypotheses. The fourth row, a single boring line of analysis, is inconsistent with the hypothesis everyone arrived believing. Note that the navigation complaints are perfectly real and still do not discriminate: people find navigation confusing in every study ever run, including the ones where usage went up.\n\nUnder step 5 you rank by evidence *against*: H2 has one inconsistency, H1 has two, H3 has two, H4 has two. H2 leads, and the next study is a composition analysis, not a usability test.\n\n## The honest part: does ACH actually debias you?\n\nMost articles about ACH stop at the matrix. The empirical record deserves to be stated plainly, because it changes how you should use the technique.\n\nACH was designed to counter confirmation bias, and the studies that have tested that claim have not supported it. Dhami, Belton and Mandel randomly assigned 50 intelligence analysts to use ACH or not on a hypothesis-testing task with probabilistic ground truth, and reported in *Applied Cognitive Psychology* (2019) that ACH-trained analysts did not follow all of the steps, that evidence for reducing confirmation bias was mixed, and that ACH may increase judgement inconsistency and error. Whitesmith's experimental work, summarized in *Cognitive Bias in Intelligence Analysis* (Edinburgh University Press, 2020), found no mitigation of confirmation bias or serial-position effects. Related work by Mandel and colleagues and by Karvetski and colleagues similarly found no advantage over an unstructured control.\n\nSo the debiasing claim is not supported. What survives is narrower and still worth your time:\n\n- **Hypothesis generation.** Being made to write down four explanations produces explanations you did not have. This is the benefit analysts consistently report, and it does not depend on the matrix scoring being correct.\n- **Evidence auditing.** The strike-out-the-non-diagnostic-rows step is a genuinely useful filter that nothing else in the standard research toolkit performs.\n- **Externalization.** The matrix makes your reasoning inspectable by someone else, which is a precondition for anyone disagreeing with you productively.\n\nTreat ACH as a thinking aid and a communication artifact, not as a bias vaccine. If your goal is specifically to stop your expectations from steering your conclusions, the intervention with better evidence behind it is [blind analysis](/docs/blind-analysis-research) - remove the information that lets you know which answer you are getting - combined with a [pre-committed analysis plan](/docs/p-hacking-researcher-degrees-of-freedom).\n\n## Common mistakes\n\n| Mistake | Why it fails | Fix |\n| --- | --- | --- |\n| Hypotheses that are not mutually exclusive | Overlapping hypotheses cannot be separated by any evidence | Rewrite so at most one can be the primary driver |\n| Only \"real\" hypotheses, no boring ones | Instrumentation, sampling and seasonality are the most common true answers | Always include at least one measurement-artifact hypothesis |\n| Scoring \"supports\" instead of \"consistent with\" | Reintroduces the confirmation logic ACH is meant to replace | Ask: if this hypothesis were true, would I expect to see this? |\n| Keeping the vivid non-diagnostic quote | It feels like evidence and moves nothing | Strike any row that is consistent across all columns |\n| Reporting only the winner | Hides how close the race was | Report the runners-up and the evidence that eliminated them |\n| Never writing step 8 | The conclusion becomes unfalsifiable | One sentence: what would change my mind |\n\n## The modern approach: making the evidence side cheap\n\nThe bottleneck in ACH has never been the matrix. It is that populating the evidence column properly requires evidence you usually do not have. You can generate four hypotheses in ten minutes and then discover that three of them are untestable with the data on hand, at which point the honest matrix is mostly blank and the team defaults back to the hypothesis they can argue for.\n\nThis is where the economics of research actually decide the analysis. Koji is built to close that gap:\n\n- **Targeted follow-up studies in days, not weeks.** Step 8 - \"what would change my mind\" - is only useful if you can go and look. AI-moderated interviews mean a hypothesis-discriminating study is a two-day exercise rather than a next-quarter commitment.\n- **[Structured questions](/docs/structured-questions-guide) produce discriminating evidence by design.** All six types - `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking`, and `yes_no` - can be written to separate hypotheses rather than confirm one. A `ranking` question that forces a choice between four candidate causes is diagnostic; an `open_ended` \"what frustrates you?\" almost never is.\n- **Automatic thematic analysis gives you prevalence, not anecdote.** Diagnosticity depends on knowing how many people said something and who did not, which is exactly what manual quote-pulling loses. See [thematic analysis](/docs/thematic-analysis-guide) and [how to analyze qualitative data](/docs/how-to-analyze-qualitative-data).\n- **Customizable AI consultants can be pointed at the rival explanation.** Running a short study designed to *break* your leading hypothesis is the ACH step nobody does, because it is a study you have to justify. When it costs a day, it stops needing justification.\n\nLegacy tooling pushes the other way. A survey platform optimized for a single dashboard encourages one hypothesis with a supporting chart. The matrix wants several hypotheses and the evidence that separates them, and that requires being able to go back and ask again.\n\n## Frequently asked questions\n\n### How many hypotheses should I list?\n\nThree to six. Below three you are not really comparing; above six the matrix becomes unreadable and the hypotheses start overlapping. If you have more, group them - and always reserve one slot for a measurement-artifact explanation, which is the single most commonly omitted and most commonly correct hypothesis in product analytics.\n\n### What is the difference between ACH and just listing alternative explanations?\n\nThe matrix and the diagnosticity filter. Listing alternatives is a brainstorm; ACH forces you to score every piece of evidence against every hypothesis and then delete the evidence that does not discriminate. The deletion step is the part that changes conclusions, and it is the part a brainstorm never reaches.\n\n### Why rank by evidence against rather than evidence for?\n\nBecause confirming evidence is abundant and usually non-diagnostic, while disconfirming evidence is rare and decisive. A single solid inconsistency can eliminate a hypothesis that a dozen consistent observations appeared to support. This is the practical form of falsification, and it is why the most probable hypothesis tends to be the one with the least evidence against it.\n\n### Does ACH work for qualitative research?\n\nYes, and arguably better than for quantitative work, because qualitative evidence is where non-diagnostic material accumulates fastest. Use coded themes as your evidence rows and theme prevalence as the consistency judgment. Pair it with [inter-rater reliability](/docs/inter-rater-reliability-qualitative-research) so the coding underneath the matrix is stable.\n\n### Is ACH proven to reduce bias?\n\nNo. Controlled studies including Dhami, Belton and Mandel (2019) and Whitesmith (2020) found mixed-to-no evidence that ACH reduces confirmation bias, and some evidence it increases judgement inconsistency. Use it for hypothesis generation, evidence auditing and making your reasoning inspectable. For debiasing specifically, blind analysis and a pre-committed plan have better support.\n\n### How long does an ACH matrix take?\n\nForty-five minutes to two hours for a real question, most of it in step 2 assembling the evidence honestly. If it takes ten minutes you almost certainly listed hypotheses that are not mutually exclusive and skipped the diagnosticity filter.\n\n## The bottom line\n\nThe value of ACH is not that it makes you unbiased - the evidence says it does not. It is that it makes the shape of your case visible: how many explanations you actually considered, which evidence separated them, and what would change your mind. Most research readouts cannot answer any of those three questions, and a one-page matrix answers all of them.\n\n## Related Resources\n\n- [Same Data, Different Answers: The Many-Analysts Problem](/docs/many-analysts-one-dataset) - how far independent analysts diverge on identical data\n- [Blind Analysis](/docs/blind-analysis-research) - the intervention with better evidence for actually reducing analyst bias\n- [P-Hacking and Researcher Degrees of Freedom](/docs/p-hacking-researcher-degrees-of-freedom) - what happens when the hypothesis is chosen after the data\n- [Confirmation Bias in User Research](/docs/confirmation-bias-user-research) - the bias ACH was designed against\n- [Structured Questions Guide](/docs/structured-questions-guide) - writing questions that discriminate between hypotheses, across all six types\n- [Conflicting Research Findings](/docs/conflicting-research-findings) - what to do when qualitative and quantitative evidence disagree\n- [Triangulation in Research](/docs/triangulation-in-research-guide) - combining sources so more of your evidence is diagnostic\n- [Research Synthesis Guide](/docs/research-synthesis-guide) - assembling findings without building a case","category":"Research Methods","lastModified":"2026-08-12T03:25:41.711317+00:00","metaTitle":"Analysis of Competing Hypotheses (ACH) for Product Research (2026)","metaDescription":"ACH scores every piece of evidence against every hypothesis to find what actually discriminates. The eight steps, a worked matrix, and the honest evidence on whether it debiases you.","keywords":["analysis of competing hypotheses","ACH method","diagnosticity","competing hypotheses matrix","rival explanations","structured analytic techniques","Heuer","hypothesis testing"],"aiSummary":"Analysis of Competing Hypotheses (ACH), developed by Richards J. Heuer, Jr. at the CIA and published in Psychology of Intelligence Analysis (1999), evaluates several explanations simultaneously using a matrix of evidence against hypotheses. Its central concept is diagnosticity: evidence consistent with every hypothesis has zero value regardless of how persuasive it feels. Rank hypotheses by evidence against them, not for them. Controlled studies (Dhami, Belton and Mandel 2019; Whitesmith 2020) found mixed-to-no evidence that ACH reduces confirmation bias, so use it for hypothesis generation, evidence auditing and externalizing reasoning rather than as a debiasing method.","aiPrerequisites":["Familiarity with forming research hypotheses","Basic experience analyzing qualitative or product data"],"aiLearningOutcomes":["Define diagnosticity and identify non-diagnostic evidence","Build an ACH matrix for a product research question","Rank hypotheses by disconfirming rather than confirming evidence","State the empirical limits of ACH as a debiasing technique","Write a falsification milestone for a research conclusion"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}