The Illusion of Explanatory Depth: Why Users Explain Confidently and Wrongly (2026)
Asking users why is asking the one question people are most overconfident about. Why mechanism probes work, reason questions do not, and what to ask instead.
The short answer
When you ask a user why they do something, you are asking the one kind of question people are most systematically overconfident about. Not mildly overconfident, and not overconfident in general: the effect is specific to explanations, and it is strongest in exactly the environments software lives in.
The finding is thirty years old and unusually well replicated. Rozenblit and Keil named it in Cognitive Science in 2002 (volume 26, issue 5, pages 521 to 562): "People feel they understand complex phenomena with far greater precision, coherence, and depth than they really do." They called it the illusion of explanatory depth, and they demonstrated it across twelve studies.
Two details in that paper change how you should run interviews, and almost nobody acts on them:
- The illusion is not general overconfidence. In their words, "The illusion is far stronger for explanatory knowledge than many other kinds of knowledge, such as that for facts, procedures or narratives." Facts, procedures and narratives are comparatively safe. Explanations are not.
- It is worst in environments like yours. "The illusion for explanatory knowledge is most robust where the environment supports real-time explanations with visible mechanisms."
A software interface is the purest example of an environment with visible mechanisms available in real time. Clicking visibly causes things. That places your users near the theoretical maximum of this illusion, and it means the confident "I use it because..." you just wrote down is the least reliable sentence in the transcript.
The good news is that there is a tested intervention, it is a single change to question form, and it is not the one most teams reach for.
The intervention: mechanism, not reasons
This is the part worth internalising, because it contradicts standard practice.
Fernbach, Rogers, Fox and Sloman published the key result in Psychological Science in 2013 (volume 24, issue 6, pages 939 to 946). Working on political policy attitudes, they found that "Asking people to explain policies in detail both undermined the illusion of explanatory depth and led to attitudes that were more moderate."
Then comes the sentence that should change your discussion guide: "Although these effects occurred when people were asked to generate a mechanistic explanation, they did not occur when people were instead asked to enumerate reasons for their policy preferences."
Read that again. Asking for a step-by-step mechanism punctured the illusion. Asking for reasons did not. Same participants, same topic, same general instruction to explain, two completely different outcomes depending on which kind of explanation was requested.
Almost every interview question in common use asks for reasons. "Why do you prefer this?" "What makes that important to you?" "Why did you choose us?" Each one is a request to enumerate reasons, which is the form the evidence says leaves the illusion intact. You get fluent, confident, coherent answers, and the fluency is the problem: it is indistinguishable from knowledge.
A mechanism request sounds different. "Walk me through exactly what happens, step by step, from the moment you open it." "Then what happens?" "And what does the system do next?" These ask the person to produce the causal chain rather than a justification for a preference. Where the chain has gaps, the gaps show, to you and to them.
One honest caveat about the import. Fernbach and colleagues studied political policy attitudes, not product usage, and I am not aware of a direct replication in a customer-research setting. What transfers with confidence is the underlying asymmetry between mechanism requests and reason requests, since that asymmetry is a property of how self-explanation works rather than of politics. Treat it as a strongly supported design principle, not as a measured effect size for your product.
Why five whys is the wrong tool here
Five whys is five consecutive requests for reasons.
That is worth stating flatly, because the technique is widely taught as a route to depth. Applied to a process you can observe, such as a manufacturing fault, it works: each answer is checkable against the machine. Applied to a person's own motivations, every iteration requests exactly the form of explanation that the evidence shows does not puncture overconfidence. You are not descending toward a cause. You are collecting five successive layers of plausible reconstruction, and each layer is generated to be consistent with the one above it.
The output feels like insight, because coherence and depth feel like understanding. That is the illusion working as documented, now with your interview structure amplifying it.
Jakob Nielsen put the practitioner version of this bluntly in a Nielsen Norman Group article on August 4, 2001: "Do not believe what people say they do."
It is worth being precise about the scope of his stronger claim, because it cuts against part of the advice below. Nielsen writes that "When talking about past behavior, users' self-reported data is typically 3 steps removed from the truth." That is a claim about retrospective reports of behaviour, which is exactly the narrative category Rozenblit and Keil found to be comparatively low in illusion.
Both can be true, and the reconciliation is the useful bit. Rozenblit and Keil are making a relative claim: narratives carry less illusion than explanations. Nielsen is making an absolute one: a remembered account of behaviour is a poor substitute for watching the behaviour. So narratives are the best thing to ask for when asking is your only option, and still worse than observation when observation is available. Ask for episodes rather than explanations, and prefer instrumentation or a recorded session over either when you can get it.
What to ask instead
Rozenblit and Keil hand you the target list directly. Facts, procedures and narratives carry far less illusion than explanations. So bias your guide toward those three.
Narratives. Ask what happened. "Tell me about the last time you did this." A specific past episode is recalled, not constructed. It has details that can corroborate or contradict each other.
Procedures. Ask what they did. "Show me how you do it." "What did you click first?" Procedural knowledge is more accurate because it is rehearsed through use rather than assembled on request.
Facts. Ask for counts and specifics. "How many times last week?" "How long did it take?" "What did you spend on this?" These are checkable and bounded.
Mechanisms, when you need causal depth. "Walk me through what the system does between you pressing save and the number updating." This is where you deliberately invoke the Fernbach effect, and where a gap is informative rather than embarrassing.
Reasons, last and with discount applied. You still want them, because a person's theory about their own behaviour is real data about how they think. Just do not record it as data about why they behave.
The ordering matters as much as the content. Ask for the narrative and the procedure first, then the mechanism, and only then the reasons. Asking why first anchors the participant to a justification they will then defend for the rest of the session, and consistency pressure will bend their subsequent factual answers toward it.
A diagnostic you can run on your last study
Open your most recent transcript set and count three things:
- Questions requesting reasons (containing why, or what makes, or how come).
- Questions requesting narrative or procedure (tell me about the last time, walk me through, show me, then what).
- Findings in your report whose only support is an answer from category one.
The third number is the interesting one. In most reports it is not zero, and frequently the headline finding sits there. A finding supported only by reason-form answers is a finding supported by the least reliable evidence class in the transcript, and it should carry a caveat or a follow-up before it drives a roadmap decision.
There is a cheap follow-up available: re-ask the same participants for the mechanism. If the causal chain they produce cannot support the reason they gave you, the reason was a reconstruction.
How Koji handles this
Consistency is the hard part. A human moderator knows all of this and still asks "why do you like it" at 4pm in their sixth session, because reason-form questions are conversationally natural and mechanism probes take effort. An AI interviewer applies the rule identically every time.
- Koji's AI interviewer probes for mechanism, on every session, without fatigue. The follow-up that matters here depends on what the participant just said, which a static survey structurally cannot do. Legacy survey tools such as SurveyMonkey can only ask the question you wrote in advance, so a reason-form question stays a reason-form question for all 400 respondents.
- The built-in Mom Test framework already encodes the right shape. Koji ships methodology frameworks including The Mom Test, Jobs to be Done, discovery, exploratory, and lead magnet research. The Mom Test question patterns are explicitly narrative and procedural: walking through how someone currently handles a task, asking what happened after that, asking about the last time something occurred, and asking what they have already tried. Those are the low-illusion forms, and the framework reaches for them by default rather than by the moderator remembering to.
- Structured questions collect the checkable classes. Koji supports six question types:
open_ended,scale,single_choice,multiple_choice,ranking, andyes_no. Use them for the fact layer. Ascaleitem for frequency, asingle_choicefor which step they actually do first, arankingfor which outcomes matter, ayes_nofor whether a thing ever happened. These sit alongsideopen_endedprobes rather than replacing them, and they are the part of your evidence that does not depend on a participant's self-theory. - Voice interviews keep narratives intact. Spoken recall of a specific episode produces more incidental detail than a text box, and incidental detail is what lets you tell a remembered episode from a constructed one.
- Analysis separates what was reported from what was explained. Because Koji grounds each extracted item in its source passage, you can check whether a finding rests on a described episode or on a participant's theory about themselves. That distinction is invisible in a spreadsheet of quotes.
- Customisable AI consultants let you enforce your own rule. If your team has decided that no finding ships on reason-form evidence alone, you can configure the interviewing behaviour to always follow a stated preference with a mechanism probe, so the rule lives in the instrument rather than in a style guide nobody rereads.
You do not need a PhD in cognitive science to apply this. You need the question forms in the right order, applied consistently across every session, which is precisely the part software does better than people.
Common mistakes
- Treating fluency as knowledge. A confident, coherent, well-structured answer about one's own motivation is the expected output of the illusion, not evidence against it.
- Asking why first. It anchors the participant to a justification they will defend, contaminating the factual answers that follow.
- Using five whys on motivation. It is five consecutive reason requests, the one form shown not to puncture overconfidence. Keep it for observable processes.
- Throwing out reason-form answers entirely. A person's self-theory is genuine data about how they think and talk, and useful for messaging. It is just not data about causes.
- Assuming a mechanism probe is rude. It is not an interrogation. "Walk me through what happens" reads as interest, and participants routinely discover their own gaps without embarrassment.
Frequently asked questions
What is the illusion of explanatory depth?
It is the documented tendency for people to believe they understand how things work with far more precision and coherence than they actually do. Rozenblit and Keil demonstrated it across twelve studies in Cognitive Science in 2002, finding that people feel they understand complex phenomena with far greater precision, coherence, and depth than they really do. Critically, it is much stronger for explanatory knowledge than for facts, procedures or narratives.
Why does this matter more for software products?
Because Rozenblit and Keil found the illusion is most robust where the environment supports real-time explanations with visible mechanisms, and a user interface is the clearest possible case of that. Clicks visibly produce effects immediately. Users therefore feel they understand the system, and by extension their own use of it, with unusually high confidence, which makes their explanations unusually unreliable.
Should I stop asking why in user interviews?
Not entirely, but demote it. Ask for narratives, procedures and facts first, since those knowledge types carry far less illusion, and use a mechanism probe when you need causal depth. Keep reason-form questions for last and treat the answers as evidence about how the person thinks rather than about why they behave as they do.
What is the difference between asking for reasons and asking for a mechanism?
A reason request asks for a justification, as in why do you prefer this. A mechanism request asks for the causal chain, as in walk me through exactly what happens step by step. Fernbach and colleagues found in Psychological Science in 2013 that mechanistic explanation requests undermined the illusion while requests to enumerate reasons did not, even with the same participants on the same topic.
Does this mean five whys does not work?
It works on observable processes, where each answer can be checked against the system. It is poorly suited to a person's own motivations, because five iterations of why are five consecutive requests for reasons, the form shown not to puncture overconfidence. The result is a deep, coherent chain of reconstruction that feels like insight because coherence feels like understanding.
How does Koji reduce this problem?
By applying the right question forms consistently, which is where human moderators drift. Koji's AI interviewer follows stated preferences with mechanism probes on every session rather than when the moderator remembers, its built-in Mom Test framework defaults to narrative and procedural patterns such as walking through a current task and asking what happened next, and its six structured question types capture the checkable fact layer alongside open-ended probing.
Related Resources
- The Mom Test for User Interviews - the methodology whose question patterns are already low-illusion
- Probing and Follow-Up Questions - the mechanics of the follow-up a mechanism probe depends on
- The Five Whys Technique in User Research - where consecutive why questions do and do not belong
- Confirmation Bias in User Research - the analyst-side counterpart to a participant-side illusion
- Thematic Analysis Guide - separating described episodes from participant self-theory during coding
- Structured Questions Guide - the six question types and which ones collect checkable facts
Related Articles
Confirmation Bias in User Research: How to Recognize and Eliminate It
Confirmation bias quietly corrupts user research by leading teams to hear what they already believe. Learn how it shows up in interviews and analysis, and the practical tactics — and AI moderation — that neutralize it.
The Five Whys Technique: How to Find Root Causes in User Research (with AI)
The Five Whys is a root-cause analysis technique that turns surface-level user feedback into actionable insight. Learn how to apply it in interviews and run it with AI-powered probing at scale.
The Mom Test: How to Ask Customer Interview Questions That Get Honest Answers
A complete guide to the Mom Test methodology by Rob Fitzpatrick—covering the three core rules, good vs. bad interview questions, avoiding confirmation bias, and how AI scales honest customer discovery conversations.
Probing and Follow-Up Questions: Going Deeper in Research Interviews
Learn the different types of probing questions — clarification, elaboration, and contrast — and when to use each to get richer qualitative data from your participants.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
The Complete Guide to Thematic Analysis
Learn how to systematically analyze qualitative data using Braun and Clarke's six-phase thematic analysis framework.