Back to docs
Research Methods

Nobody Made the Decision: How to Research a Purchase No Single Person Chose (2026)

58% of CEOs' retrospective accounts of their own firm's strategy disagreed with their own earlier validated reports. Why 'why did you buy' is unanswerable for group decisions - and what to measure instead.

You can fix informant selection. You can eliminate proxy bias. You can run a flawless interview with a perfectly chosen participant who remembers everything - and still fail to recover why an account bought or left, because the answer was never located in any single person's head to begin with.

Here is the finding that should reset expectations. Golden asked chief executives to report their firm's current strategy. Two years later, he asked the same executives what their firm's strategy had been two years earlier - the exact period they had already described. 58% of those retrospective accounts did not agree with their own previous, validated reports (Golden, Academy of Management Journal, 1992, 35(4), 848-860). And the errors were not random: they "appear to occur systematically and may be attributable to faulty memory or to attempts to cast past behaviors in a positive light."

If a CEO cannot reliably recall their own company's strategy, no informant-selection rule is going to recover "why we chose you" from a buying group of nine people, eleven months after the fact.

This is a different kind of problem from the previous two

This is the third article in a chain, and it is the one that does not have a fix in the same shape as the others.

  • Key informant error is a selection problem. Choose better informants, use more than one, and the error shrinks.
  • Proxy response bias is a routing problem. Send each question to someone who could have observed the answer, and the error shrinks.
  • This is neither. It is a problem with the thing you are trying to measure. There is no single reason a company bought, in the same sense that there is no single reason a parliament passed a law. Improving your instrument does not help, because the quantity you are pointing it at does not exist in the form your question assumes.

That is why it survives fixing the other two. An organization that has solved informant selection and proxy routing still writes "the deal was lost on price" in its win-loss deck, and that sentence is still an artifact of the method rather than a fact about the world.

Three reasons the decision is not in anyone's head

1. It was distributed. Nobody saw all of it.

A B2B purchase is assembled over months out of a security review nobody outside security watched, a budget conversation that happened without the champion, a bake-off run by one team, and a procurement negotiation that changed the terms after the technical decision was made. Every participant saw a slice. Asking any one of them for "the reason" asks them to narrate the parts they did not attend.

The size of that effect is measurable. In a study of eight multi-informant datasets, 33.7% of the construct-level variance in intraorganizational measures was attributable to which type of informant was asked (Homburg, Klarmann, Reimann and Schilke, Journal of Marketing Research, 2012, 49(4), 594-608). A third of what you hear about the inside of a company is the shape of the seat the person was sitting in.

2. The outcome rewrote the memory.

Everyone you interview already knows how it turned out. That knowledge is not neutral. In their review of the field, Roese and Vohs (Perspectives on Psychological Science, 2012, 7(5), 411-426) describe hindsight bias as running on three engines at once: cognitive (people selectively recall information consistent with what they now know to be true, and engage in sensemaking to impose meaning on their own knowledge), metacognitive (the ease of understanding an outcome gets misread as evidence it was always likely), and motivational (a need to see the world as orderly and predictable, and to avoid being blamed for problems).

The consequence they name is precisely the failure mode of a win-loss program: "myopic attention to a single causal understanding of the past (to the neglect of other reasonable explanations) as well as general overconfidence in the certainty of one's judgments." Your participants are not withholding the reason. They have been handed a tidy one by their own cognition, and they believe it.

3. Every account is role-shaped, and the roles have stakes.

The champion who lost the internal argument has a reason to describe the decision as irrational. The economic buyer has a reason to describe it as disciplined. The person who ran the evaluation has a reason to describe it as thorough. None of them is lying. All of them are producing an account that is coherent from where they sat and from what it costs them to be wrong.

What this does not mean

Two clarifications, because the argument is easy to over-extend.

This is not memory decay. Recall bias covers the way memory degrades and distorts over time, and everything in that guide still applies. But this problem persists at zero elapsed time: interview all nine participants the afternoon the contract is signed and you still will not have "the reason," because a perfect memory of your own part does not contain the whole. Recency helps with decay. It does nothing for distribution.

Retrospective research is not worthless. The strongest counterweight to Golden's finding came five years later: Miller, Cardinal and Glick's reexamination (Academy of Management Journal, 1997, 40, 189-204) argued the situation was not as dire as claimed, and concluded that retrospective reporting is a viable methodology when the measure used to generate the reports is adequately reliable and valid - to be neither rejected nor used indiscriminately. That conclusion is the hinge of this whole article. It does not say "ask better people." It says the measure has to change.

Change the estimand: stop asking for reasons, start reconstructing events

The single move that resolves this is to stop treating "why" as the thing being measured.

"Why did you choose us?" asks a participant to perform a causal analysis of a multi-person process using only their own memory - which is the one task the evidence above says they cannot do. "What happened on the 14th?" asks them to report an event they attended. The first produces a plausible story. The second produces a data point that can be corroborated, contradicted, and assembled with other people's data points into something none of them individually possessed.

The estimand shifts from the reason (which does not exist as a single quantity) to the decision process (which does, and leaves traces).

The four-part protocol

1. Elicit events, not reasons. Every question should be answerable by someone who was in the room, about something dated and witnessed.

Instead ofAsk
"Why did you choose the other vendor?""What was the last meeting where both vendors were still options? Who was in it?"
"What was most important in the decision?""What changed between the shortlist and the final call?"
"Was price the issue?""When did price first come up, and who raised it?"
"Who was the decision maker?""Who could have stopped this, at which stage?"

The last one is worth its own note. "Who decided?" invites a tidy answer; "who could have stopped it, and when?" maps the real veto structure, which is what actually determined the outcome.

2. Elicit artifacts, not just accounts. Organizational decisions leave a paper trail that does not rewrite itself after the fact: the business case, the evaluation scorecard, the security questionnaire, the procurement checklist, the Slack thread, the calendar. Asking "can you walk me through the scorecard you used?" gets you a contemporaneous record instead of a reconstructed one. This is the practical version of Miller, Cardinal and Glick's condition - it improves the measure, not the respondent.

3. Force alternative explanations. Roese and Vohs end their review with the one intervention that has been shown to work: encouraging people to consider alternative causal explanations for a given outcome reduces hindsight bias. That is directly operationalisable as a probe, and it belongs after every causal claim a participant makes:

"You have said the price was too high. Suppose price had been identical - what else, if anything, would have gone differently?"

"What is the strongest case someone on your team could make that this was decided for a completely different reason?"

Participants who confidently name one cause will, under this probe, frequently name two more and downgrade the first. That is not the interview failing. That is the interview working.

4. Treat disagreement as the finding, not the error. When your five participants give you three incompatible accounts, the temptation is to reconcile them into one narrative or to weight the most senior one. Both destroy the actual result. The correct output is a divergence map: which claims all participants agreed on (those are close to facts), which split cleanly by role (those tell you where the internal conflict was), and which only one person made (those are hypotheses).

An account where the champion, the end user and the finance approver tell you three different stories is not a failed study. It is a finding about that account - and usually a more useful one than the consensus you would have manufactured.

What to report instead of a reason code

A win-loss or churn readout built this way looks different, and is more defensible:

Conventional outputEvent-based output
"Lost on price (38% of losses)""In 12 of 31 losses, a budget owner entered the process after the technical shortlist was set"
"Churned due to lack of adoption""In 9 of 14 churns, the original champion had left the account before renewal"
"They wanted feature X""Feature X was named in the RFP in 6 cases; in 4 of those it was not mentioned again after the demo"

The right-hand column is made of things participants observed and artifacts corroborated. It also tells you what to change, which the left-hand column never does - "lost on price" has been the top reason code in every deck since decks existed, and no team has ever known what to do with it.

The modern approach: how Koji makes this practical

Reconstructing a distributed decision means interviewing four to nine people per account, quickly, before the accounts converge - and doing it consistently across dozens of accounts. That is why almost nobody does it: as a manual program it is arithmetically impossible.

  • Parallel, asynchronous interviews stop accounts contaminating each other. When interviews are scheduled sequentially over three weeks, participants talk to each other in between, and the organization's official story hardens. Koji's async AI interviews go out to every participant at once and come back within days, capturing accounts before they converge on the tidy version.
  • The same instrument for every participant. Role-shaped differences are only interpretable if the questioning did not vary. A single AI interviewer removes the moderator variance that would otherwise be tangled up with the divergence you are trying to measure - the problem covered in the AI interviewer house effect guide.
  • The alternative-explanation probe fires every time. This is the de-biasing intervention with the best evidence behind it, and it is also the probe a human moderator most often skips because it feels confrontational. An AI interviewer runs it after every causal claim, in every interview, without fatigue and without social cost to the participant.
  • Structured questions make divergence measurable. All six types earn their place here: ranking to have each participant order the same set of factors so you can compute how far apart they are, single_choice for who had veto authority at each stage, scale for confidence in their own account, multiple_choice for which artifacts existed, yes_no for whether a given meeting happened, and open_ended for the event narrative that the AI then probes. Divergence stops being an impression in a synthesis session and becomes a number.
  • Timeline assembly across transcripts. Because every interview is transcribed and analysed automatically, dated events from nine separate accounts can be merged into one chronology, with the contradictions preserved rather than smoothed. That is the artifact no reason-code survey can produce.
  • Real-time reporting while the trail is warm. Themes and quotes appear as interviews complete, so a second wave of questions can chase a contradiction the first wave exposed - within the same week, while the artifacts are still findable.

Traditional win-loss surveys ask a departing buyer to pick their primary reason from a list. That instrument cannot fail to produce a reason, whether or not one exists - which is precisely why its output has never been actionable.

Frequently asked questions

Why can't customers tell me why they bought?

Individual customers often can tell you a great deal about their own part. What no individual can supply is the reason a group decided, because organizational decisions are distributed across people and months, and each participant witnessed only a slice. Their account is then reshaped by knowing the outcome - hindsight bias produces confident, tidy, single-cause explanations that neglect other reasonable ones.

How reliable are retrospective accounts of business decisions?

Poor, when the measure is a global judgement. Golden (1992) found that 58% of chief executives' retrospective accounts of their own firm's earlier strategy disagreed with their own previous validated reports. Miller, Cardinal and Glick (1997) later argued retrospective reporting is still viable - but only when the measure generating the reports is adequately reliable and valid, which points to event-level questions and documentary artifacts rather than "why" questions.

Is this just recall bias?

No. Recall bias is about memory degrading over time, and interviewing sooner reduces it. This problem is present even at zero delay: a participant with a perfect memory of their own involvement still does not hold the parts of the decision they never saw. Speed helps with one and not the other.

What should I ask instead of "why did you choose us?"

Ask about dated, witnessed events: what was the last meeting where more than one option was live, what changed between the shortlist and the final decision, when did price first come up and who raised it, and who could have stopped the purchase at each stage. Then ask for the artifacts - the business case, the scorecard, the RFP - which record what was thought at the time rather than what is remembered now.

How many people do I need to interview per account?

Enough to cover the distinct vantage points rather than a fixed headcount: typically the champion, an end user, the economic approver, and whoever ran the evaluation. Interviewing four people who all sat in the same function is close to interviewing one. The test is not how many participants you have but how many independent vantage points on the decision they represent.

What do I do when participants contradict each other?

Report the contradiction. Claims every participant agrees on are close to facts; claims that split cleanly along role lines tell you where the internal conflict was and are usually the most commercially useful output; claims only one person makes are hypotheses for the next study. Averaging or deferring to the most senior participant destroys the most informative part of the dataset.

Related Resources

Related Articles

Recall Bias: How Faulty Memory Distorts Research (and How to Prevent It)

Recall bias is the systematic error that arises when respondents remember past events inaccurately or incompletely. Learn why memory is reconstructed not retrieved, how telescoping distorts data, and how to design around it.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Switch Interviews: The JTBD Method for Understanding Why Customers Buy (and Leave)

Switch interviews uncover the four forces of progress that cause customers to switch from one product to another. Learn the Bob Moesta playbook and how to run switch interviews with AI at scale.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.

The 5-Second Test: How to Measure First Impressions and Visual Hierarchy (2026 Guide)

A complete guide to the 5-second test — the lightweight UX research method that measures gut reactions, message clarity, and visual hierarchy. Learn how to design questions, recruit participants, analyze results, and combine 5-second tests with AI interviews.

A/B Testing vs. User Research: When to Use Each (And When to Use Both)

Understand when A/B testing and qualitative user research each shine, and how to combine them for better product decisions. Includes framework for choosing methods, real case studies, and how AI interviews make mixed methods accessible.