Back to docs
Research Operations

Outcome Bias: Why You Cannot Grade a Decision by Its Result (2026)

Decision quality is never observed; only outcomes are, and an outcome is decision quality plus luck. Why this has no technical fix, and what you can audit instead.

Answer first: you cannot grade a decision by how it turned out, because the thing you want to grade - decision quality - is never observed. Only outcomes are, and an outcome is decision quality plus luck. Show people identical decisions with different endings and they rate the lucky one as better, by roughly a full point on a seven-point scale, with an effect size above 1.0. They do it even when they have just said outcomes should not matter. This is the last and hardest of the measurement problems, because unlike sampling error or wording effects, it has no technical fix: the counterfactual you would need does not exist. The only auditable object is the decision record you wrote before you knew.

The experiment that established it

Jonathan Baron and John Hershey published "Outcome bias in decision evaluation" in the Journal of Personality and Social Psychology (54(4):569-579, 1988). The design is almost annoyingly simple, which is why it has held up.

Participants read about a 55-year-old man with a heart condition considering a bypass operation. The operation would relieve his pain and raise his life expectancy from about 65 to about 70. They were told that "8% of the people who have this operation die from the operation itself." They then rated the decision to operate on a seven-point scale running from 3, "clearly correct and the opposite decision would be inexcusable," to -3, "incorrect and inexcusable."

Every participant saw the same information the decision maker had. The only thing that varied was how it ended: the patient recovered, or the patient died. Baron and Hershey found that the decision was rated better when the outcome was good - effect sizes of about d = 0.21 when the patient decided and d = 0.53 when the physician did.

The result has since been re-run at modern sample sizes. A registered replication of Experiment 1 published in the International Review of Social Psychology (2022) recruited 692 participants and found the effect substantially larger than the original:

ConditionOutcomenMean ratingSD
Physician decidedSurvived1731.810.84
Physician decidedDied1720.451.55
Patient decidedSurvived1711.760.79
Patient decidedDied1760.901.36

Effect sizes were d = 1.10 in the physician condition and d = 0.77 in the patient condition. The authors described it as "a successful replication with signal and direction consistent with that of the target article's findings."

Then the finding that should end the argument. A subgroup of participants (n = 44) separately reported that outcomes "definitely" or "probably" should not influence how a decision is evaluated. They showed the bias anyway: d = 0.64, p = .03.

Knowing about outcome bias does not protect you from outcome bias. That is what makes it a structural problem rather than a training problem.

The capstone claim: the quantity you want is never observed

This library contains a long series of articles about things that make a measurement wrong. Total survey error prices seven of them against a fixed budget. Sampling bias is about who got in. Statistical power is about whether the instrument could have seen the effect. Every one of them shares a structure: there is a true value out there, your estimate misses it, and the article tells you how to miss by less.

Decision quality does not have that structure, and this is the point of this article.

Decision quality is not a hard-to-measure quantity. It is an unobservable one. A good decision is one that was correct given what could be known at the time, over the distribution of things that might have happened. Exactly one of those things did happen. The others - the ones that define whether the decision was right - did not, and never will. There is no larger sample, no better instrument and no cleaner design that recovers them, because the counterfactual is not missing data. It does not exist.

What you observe instead is the outcome, which is:

Outcome = decision quality + luck

Two unknowns, one equation, one observation. No amount of rigour solves an underdetermined system. This is why outcome bias is not a lapse in judgement that better-trained reviewers avoid: reviewers are asked to report a quantity they cannot see, and they substitute the only correlated quantity available.

That substitution is not irrational, either. Over many decisions, outcome really is evidence about quality - the luck term averages out. The error is applying that logic to n = 1, where it does not.

Why this is not fixed by measuring harder

Three properties make outcome bias resistant to the standard remedies.

It survives knowing about it. The replication's d = 0.64 among people who had just endorsed the correct principle is the cleanest demonstration available. Compare this with most measurement problems, where naming the mechanism largely defuses it.

It is asymmetric in organisations. Bad outcomes get reviewed; good ones do not. A launch that succeeds is never subjected to a post-mortem asking whether the reasoning was sound, so bad reasoning that got lucky is never corrected and enters the culture as a precedent. This is the survivorship mechanism applied to decisions rather than customers, and it means the correction is systematically one-sided.

It has second-order effects on behaviour. People who expect to be judged on outcomes take decisions that produce defensible outcomes rather than good ones. The characteristic symptom is risk aversion on exactly the high-variance bets that portfolio evidence says carry the returns - which is where this connects to the base rate for a product bet. If two-thirds of well-reasoned ideas fail, and failure is punished at the individual level, the rational individual response is to stop proposing ideas that could fail visibly. The organisation gets a safer, worse portfolio, and no single person did anything unreasonable.

Reconciling this with calibration scoring

There is an apparent contradiction with calibration scoring for research teams, which argues for grading research claims against what actually happened using proper scoring rules. If outcomes cannot grade decisions, how can outcomes grade forecasts?

The reconciliation is about n, and it is worth stating precisely because it determines which tool you reach for.

Grading one decisionGrading a track record
What you observeOne outcomeMany outcomes
The luck termDominates; cannot be separatedAverages toward zero
Valid methodReview the process and the recordProper scoring rules across all claims
Invalid method"It worked, so it was a good call""Their last call was wrong, so they are miscalibrated"

A Brier score over fifty claims is informative precisely because the noise cancels. A Brier score over one claim is a coin flip with a decimal point. So the rule is: score the record, review the decision. Both articles are right within their n, and the failure mode in each direction is the other one's method applied at the wrong scale.

This also explains why a single spectacular success or failure is such poor evidence about a person or a team, and why organisations that promote on the strength of one shipped hit are running the invalid method in the left-hand column.

What you can grade: the decision record

If the outcome cannot be the object of review, something else has to be, and it has to be created before the outcome is known. That artefact is a decision record, and it is the only defence with any evidence behind it.

The fields that make a decision reviewable:

FieldWhat it capturesWhy it is needed later
The decision and its ownerWhat was decided, by whom, whenPrevents retrospective reassignment of authorship
The options consideredThe two or three real alternativesA decision with one option was not a decision
What was knownThe evidence available, with linksStops hindsight importing facts that arrived later
What was uncertainThe things nobody knew, stated plainlyThe luck term, named in advance
The expected outcomeA probability or a range, not a hopeMakes the claim scoreable across many decisions
The switch pointWhat would have changed the callTies back to the value of information
The review dateWhen this gets looked at againTurns a document into a loop

The discipline that gives this its power is that the record is written before the outcome. Hindsight reliably rewrites what people believe they knew, so a record reconstructed after the fact is not evidence about the decision - it is evidence about the reconstruction.

The four questions of an honest decision review

With a record in hand, a review can ask questions that have answers:

  1. Given only what is in the record, was this the best available option? Not "did it work." Ban outcome information from the first pass entirely - the same logic that makes blind analysis work.
  2. Was anything knowable at the time missing from the record? This is the real learning question, and it is the one that improves future decisions. A study that could have been run cheaply and was not is a genuine process failure; an unknowable fact is not.
  3. Did the outcome fall inside the stated range? If a range was written down and the result landed outside it, the model of the world was wrong, and that is separable from the decision being wrong.
  4. Only now: what did the outcome cost or earn, and does anything need repairing? A real question, but a separate one from decision quality, and running it fourth prevents it colouring the first three.

A review that starts at question four - which is what almost every post-mortem does, because that is what triggered the meeting - will produce outcome bias reliably. The order is the intervention.

How Koji supports this

The hard part of the decision record is not the template. It is that the evidence it points at has to exist, be linked, and still be readable a year later when the review happens.

  • Studies stay attached to the decision. A Koji report is a durable, shareable link rather than a deck that gets edited between meetings, so "what was known" in a decision record points at the actual instrument, sample and result rather than a summary of them. See publishing and sharing reports.
  • Structured questions make expectations scoreable. Six question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - turn "customers will like this" into a distribution on a defined scale. A stated expectation of "median 7 or better on the 0-10 scale" can be checked at review time; "we thought it would land well" cannot.
  • Speed lets you record the expectation before you commit. The reason expectations go unrecorded is usually that the decision moves faster than the research. When a study fields and reports in a day - Koji runs the interviews itself in voice or text and writes its own follow-up questions from each answer - the record can be written while the decision is still open rather than reconstructed after it.
  • The AI moderator does not know the outcome. Any follow-up research after a launch is conducted by a moderator with no stake in whether the launch is judged a success, which removes one channel through which a known outcome shapes the evidence collected about it.

Common mistakes

  1. Reviewing only failures. The asymmetry is the mechanism. If you only review bad outcomes, you only ever correct the reasoning that got unlucky, and lucky bad reasoning becomes precedent. Sample some successes.
  2. Writing the decision record after the outcome. This is not a record, it is a reconstruction, and hindsight has already edited it.
  3. Judging a person on one call. Valid at large n, invalid at n = 1. See the table above.
  4. Confusing a bad outcome with a bad decision, or a good outcome with a good one. Both errors are the same error. The second is more dangerous because nobody complains about it.
  5. Recording a hope instead of an expectation. "We expect this to do well" cannot be scored. "We expect 20-30% of trials to convert, and would be surprised outside 15-40%" can.
  6. Assuming awareness is enough. The replication participants who said outcomes should not matter still showed d = 0.64. Structure beats intention here; the four-question order is the structure.

Frequently asked questions

What is outcome bias?

Outcome bias is the tendency to judge the quality of a decision by how it turned out, even when the person judging has exactly the information the decision maker had. Baron and Hershey demonstrated it in 1988 using a medical scenario in which only the ending varied, and a 2022 replication with 692 participants found effect sizes of d = 0.77 to d = 1.10 - a decision rated around 1.8 on a seven-point scale when it worked and as low as 0.45 when the identical decision failed.

How is outcome bias different from hindsight bias?

Hindsight bias is about knowledge: once you know what happened, you misremember having predicted it. Outcome bias is about evaluation: you judge the decision more harshly or more kindly because of the result, even while correctly recalling the information available beforehand. They usually appear together and reinforce each other, but the remedies differ - hindsight bias is countered by a written record of what was believed, outcome bias by reviewing the decision before the outcome is disclosed.

Can we not just train people to ignore outcomes?

The evidence says no, at least not by awareness alone. In the 2022 replication, participants who explicitly reported that outcomes definitely or probably should not influence evaluation still showed the bias at d = 0.64. Structural remedies work better than intentions: write the decision record before the result is known, review the reasoning before disclosing the outcome, and sample successes for review rather than only failures.

If outcomes cannot grade decisions, how can Brier scores grade forecasts?

The difference is sample size. Over many claims, the luck component of outcomes averages toward zero and what remains is signal about judgement, which is exactly what a proper scoring rule extracts. Over a single claim, luck dominates and no scoring rule can separate it from quality. The working rule is to score the track record and review the individual decision, and to treat any confident judgement of one decision from its outcome as unsupported.

What goes in a decision record?

The decision and its owner, the options genuinely considered, what was known with links to the evidence, what was uncertain, an expected outcome stated as a probability or range, the result that would have changed the call, and a date for review. The single most important property is that it is written before the outcome is known - a record reconstructed afterwards documents the reconstruction rather than the decision.

Does this mean we should stop holding teams accountable for results?

No. It means separating two accountabilities that usually get merged. Teams remain accountable for outcomes at the portfolio level, where luck averages out and the numbers mean something. Individual decisions are held accountable for process: were the options considered, was cheap available evidence gathered, was the expectation stated in advance. Merging the two produces predictable risk aversion, which is expensive in a domain where roughly one idea in three works and the returns sit in the tail.

Related Resources

Related Articles

Blind Analysis: How to Analyze Research Before You Know the Answer

Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.

The Base Rate for a Product Bet: What Experiment Portfolios Say About How Often a Feature Works

Roughly one product idea in three improves the metric it was built for. The published portfolio numbers, how to build your own reference class, and why idea variance beats sample size.

Evidence or Ammunition: What a Study Is For When No Result Could Change the Decision

A study commissioned after the decision is settled is not waste, it is a signal. How to tell whether you are being asked for evidence, input or support, before fieldwork starts.

Calibration Scoring for Research Teams: How to Find Out If Your Insights Were Actually Right (2026)

Research is graded on process and almost never on outcome. Forecasting tournaments solved this with proper scoring rules. Here is how to score a research team on whether its claims came true.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Survivorship Bias in Customer Research: Why You're Only Hearing Half the Story

Survivorship bias makes customer research dangerously optimistic by only sampling the customers who stayed. Learn how to spot it, why it inflates every metric, and how to systematically capture the voices of the customers who left.

Total Survey Error: The Seven Ways a Study Is Wrong (and How to Spend a Fixed Budget Across Them)

Sample size buys down exactly one of seven error components. Learn the total survey error framework, why federal agencies report only the computable one, and how to write a one-page error budget before you field.

Expected Value of Information: How to Decide Whether a Study Is Worth Running (2026)

If no realistic result would change what you do, the value of the information is zero and the correct sample size is zero. How to run the value-of-information test before you plan a study.