When No Study Was Wrong: Why Research Programs Fail Without a Defective Study
Some of the worst research-driven decisions trace to no bad study at all. Every study was true; the loss came from the interactions between them. The safety-engineering framework for losses with no component failure, applied to research.
Bottom line up front: Some of the worst research-driven decisions a company makes cannot be traced to a bad study. Every study was well designed, well run and true. The loss came from the interaction between them: a finding carried outside the population it was measured on, two true claims combined into a false one, a constraint that used to hold and quietly stopped. Safety engineering names this class - a system accident, a loss with no component failure - and it has a method for it that root cause analysis structurally cannot supply. If your only repair mechanism is finding a defective part, you will keep finding a defective person instead.
The failure with no defective part
Nancy Leveson, professor of aeronautics and astronautics at MIT, describes a chemical plant loss in Safety Science 42(4):237-270, 2004. A computer controlling a reactor was told not to change anything while an alarm was unhandled; an alarm came in; the computer therefore held the cooling-water flow at a low rate; the reactor overheated and vented to the atmosphere. Her comment is the sentence this entire article exists to apply:
"Note that there were no component failures involved in this accident: the individual components, including the software, worked as specified but together they created a hazardous system state. Merely increasing the reliability of the individual components or protecting against their failure would not have prevented the loss."
Charles Perrow coined the term system accident for this: a loss that "arises in the interactions among components (electromechanical, digital, and human) rather than in the failure of individual components." Leveson contrasts it with the ordinary kind, the component failure accident, and notes that the interaction class "received less attention" precisely because in simpler systems you could test your way out of it.
Product research organisations are now firmly in the first category and are still equipped only for the second. The two preceding articles in this series gave you component-level tools: pipeline yield tells you which stage is weakest, and the all-or-none inspection rule tells you whether checking that stage is worth it. Both assume the loss is inside a stage. This article is about what is left when both are fixed, and it is not a residue. It is most of the interesting failures.
You will always find a defendant, and that is the problem
Here is the trap. When a decision goes wrong, the organisation runs a retrospective, and the retrospective finds something. It always finds something. Leveson explains why, and the mechanism transfers to research without modification:
"following an accident, it will be easy to find someone involved in the dynamic flow of events that has violated a formal rule by following established practice rather than specified practice."
Real work always deviates from written procedure, because written procedure never covers the actual conditions. So a search for a rule violation always succeeds. The published consequence is stark: "it is not surprising that operator 'error' is found to be the cause of 70-80% of accidents." That figure is not a finding about operators. It is a finding about where investigations stop looking.
Translate it. After a product failure, someone will discover that the researcher recruited from the existing customer list rather than the specified frame, or that the readout compressed a hedged finding into a headline, or that the PM extended a result about enterprise admins to self-serve users. All of those are true. None of them is the cause. Each of them is normal practice that was fine in every other study, and each one gives the organisation the thing it is looking for - a defective part - which lets it close the review without changing anything about how the parts interact.
The tell that you are in this class: the fix that comes out of the retrospective is a reminder. Remind researchers to state their sample frame. Remind PMs to read the caveats. Reminders are what an organisation produces when it has correctly sensed there was no defect and still needs to name one.
The losses live at the boundaries
If the failure is not inside the parts, it is between them, and there is a measurement of that. Leveson cites Jacques Leplat's study of the steel industry, which "found that 67 percent of technical incidents with material damage occurred in areas of co-activity, although these represented only a small percentage of the total activity areas."
Two-thirds of the damage happened in the small fraction of space where two activities overlapped. Leveson generalises the mechanism: "Overlap areas exist when a function is achieved by the cooperation of two controllers or when two controllers exert influence on the same object. Such overlap creates the potential for conflicting control actions."
A research function is dense with overlap areas and nobody maps them. The insights repository is written by researchers and read by product managers who were not in the room. A segmentation is owned by marketing and used by product. A tracker is fielded by an agency and interpreted in-house. A finding about churn drivers is produced by a research team and applied by a growth team to a different cohort. Every one of those is a co-activity area where two parties exert influence on the same object, and none of them is anyone's job.
In the 1994 Black Hawk shootdown that Leveson analyses, one of the structural findings was that "an Army base controlled the flights of the Black Hawks while an Air Force base controlled all the other components of the airspace," and that "a common control point once again was high above where the accident occurred in the control structure." That is the shape to look for: the first person with authority over both sides of the interaction is three levels up and was never consulted.
Four ways true findings combine into a false conclusion
| Pattern | What happens | The tell |
|---|---|---|
| Scope transfer | A finding measured on one population is applied to another. Both the finding and the application are defensible in isolation | The claim in the deck has no population attached; the study had one |
| Conjunction | Two true findings from separate studies are read together to imply a third claim that neither supports and nobody tested | The synthesis slide has two citations and three claims |
| Stale constraint | A finding was true and the condition that made it true has since changed. No study was wrong; the world moved | Nobody can say what would have to change for the finding to stop holding |
| The question nobody owned | The decisive question falls between two teams' remits, so every study is well run and the question is never asked | Every study passes review and the decision still feels unsupported |
None of these is detectable by auditing a single study, which is why they survive a yield audit and an inspection regime. Each of them is a property of the set of studies and how they were used, and the set has no owner.
Scope transfer is the most common and the most quantifiable. It is worth reading alongside survivorship bias in customer research and sampling bias, which explain how a population gets silently swapped. The stale-constraint pattern is the one research refresh cadence addresses directly, with half-lives by finding type.
Treat it as a control problem, not a failure problem
The methodological pivot is the useful part. The STPA Handbook, written by Leveson and John Thomas at MIT and published free in 2018, states the underlying model:
"In addition to component failures, STPA assumes that accidents can also be caused by unsafe interactions of system components, none of which may have failed."
And the consequence for what you do about it:
"In STAMP, safety is treated as a dynamic control problem rather than a failure prevention problem. No causes are omitted from the STAMP model, but more are included and the emphasis changes from preventing failures to enforcing constraints on system behavior."
The handbook is equally clear about why the tools most organizations already have do not reach this class. "Traditional hazard analysis techniques such as FMEA and fault tree analysis focus on individual component failure or faulted modes. This focus on component reliability assumes that robust partitioning and interface control are valid and that components do not interact either directly or indirectly. However ... these assumptions usually do not hold for complex systems." Considering causes one at a time, it notes, "essentially reduces to a FMEA where only single component failures are considered."
The property being controlled is what systems theory calls emergent: something that is "not in the summation of the individual components but 'emerge[s]' when the components interact." Decision quality in a research organisation is exactly that. No single study contains it. It exists in how studies are combined, scoped and carried.
The research control structure
Rewriting the model for a research function gives you something you can actually build: a short list of constraints, and for each one, a named controller and the feedback that tells them whether it is holding.
| Constraint on the system | Who enforces it | Feedback that proves it holds |
|---|---|---|
| A finding is never stated without the population it was measured on | Whoever publishes to the repository | Sample every claim in the current roadmap deck; count those with a population attached |
| A claim that combines two studies is labelled as an inference, not a finding | Synthesis author | Count multi-source claims that name their inference step |
| Every live finding carries an expiry condition, not just a date | Finding owner | Fraction of repository entries that state what would falsify them |
| Questions that fall between two teams are assigned, not dropped | The first common manager of both teams | A standing list of unowned questions, reviewed monthly |
| A decision records which findings it relied on | Decision-maker | Traceability of roadmap items back to evidence |
Note what these are not. None of them is a quality bar on an individual study. Every one of them is a constraint on the interaction between studies, or between a study and its use. That is the whole move.
The last row is the one most teams can start on tomorrow, and there is a working procedure for it in the roadmap evidence audit. The fourth row - the unowned question - is the hardest and the highest value, and it connects directly to group decision reconstruction: if no single person made the decision, no single person noticed the question was missing either.
Why root cause analysis finds nothing here
Root cause analysis is a genuinely good tool and it is the wrong tool for this class, for a structural reason rather than a quality reason. Event-chain methods work backwards from a loss through a sequence of events until they reach something that can be called a cause and fixed. That procedure requires the chain to exist. In a system accident the components did what they were specified to do, so there is no anomalous event to find; the chain runs backwards through a series of correct behaviours and terminates at whoever the investigator gets tired of exonerating.
Leveson's summary of why local reasoning cannot get there is the best single sentence in the literature for anyone running a research organisation:
"It is difficult if not impossible for any individual to judge the safety of their decisions when it is dependent on the decisions made by other people in other departments and organizations."
Substitute "the validity of their conclusions" for "the safety of their decisions" and you have described every synthesis meeting in a company with more than one research team.
Running the review that does work
When a decision goes wrong and every study checks out, run this instead of a root cause analysis.
- Write the loss, not the error. "We built a self-serve onboarding flow that nobody used," not "the research was wrong."
- List every finding that fed the decision, with its population and its date. Do not evaluate them. Most will be fine.
- Mark the boundaries. For each finding, write who produced it and who used it. Any row where those differ is a co-activity area.
- Test each boundary for the four patterns. Was the population transferred? Were two findings conjoined? Had a condition changed? Was there a question that belonged to neither party?
- Find the constraint that was not enforced, and the controller who could have enforced it. There is always one, and it is almost never the researcher or the PM. It is usually the first person with authority over both sides.
- Fix the feedback, not the person. If the constraint has no feedback signal, the controller cannot enforce it, and adding a reminder will not create one.
The output of this review is a constraint with a named owner and an observable signal. The output of the review it replaces is a reminder.
How Koji helps, and where it does not
Tooling cannot enforce a constraint nobody has written down. What it can do is make the feedback cheap enough that the constraint is enforceable at all.
- Population travels with the finding. In Koji, every claim in a report is generated from interviews that carry their own study, screener and participant metadata, so the population is attached to the finding rather than living in a separate methods slide that does not get copied into the deck. Scope transfer becomes visible instead of silent.
- Stable question IDs make conjunction auditable. Koji's six structured question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - carry stable IDs from interview plan through analysis to report. When two studies are combined, you can tell whether they actually asked the same question or two questions that sound alike.
- Re-asking is cheap enough to test a stale constraint. The stale-constraint pattern persists because re-fielding costs weeks. When a study can be re-run against fresh participants in hours with an AI moderator, "what would have to change for this to stop being true" becomes a question you can answer rather than a caveat you write.
- The unowned question is still yours. No platform will notice that the decisive question falls between two teams. That is a control-structure problem and it is fixed by assigning a controller, which is a management act. Koji shortens the loop once someone has decided to ask.
The honest framing: legacy research tooling is built around the study as the unit, which is precisely the component-level assumption that makes this failure class invisible. An AI-native platform where studies are cheap and their metadata is machine-readable at least makes the interactions inspectable. It does not make them managed.
Honest objections
"This is an excuse for bad research." It is the opposite. It only applies once the component-level work is done. If your screening yield is 19 percent, you do not have a system accident, you have a defective part, and you should go and read the yield article first. The capstone class is what remains after that.
"Safety engineering is a different domain with different stakes." The stakes are different and the causal structure is not. Leveson's argument is about complexity and interaction, not about hazard severity, and the tools she is criticising - fault trees, FMEA, five whys - are the same tools product organisations imported.
"If nobody is at fault, nothing gets fixed." That is the failure mode this replaces, not the one it creates. The STAMP answer is that responsibility attaches to whoever owned the constraint that was not enforced. That is a real accountability claim and it is usually more senior than the person a root cause analysis would have named.
"We do not have a control structure to draw." You do; it is undrawn. The five-row table above is a starting version. Write who enforces each constraint. The rows where the answer is "nobody" are your finding.
Frequently asked questions
What is a system accident, and how is it different from a mistake?
A system accident is a loss that arises from the interactions among components rather than from any component failing. Leveson's chemical-reactor example is the canonical case: the software behaved exactly as specified, and the specified behaviour combined with an alarm condition to overheat the reactor. Applied to research, it is a wrong decision reached from studies that were all well run and all true.
Why does root cause analysis fail on this class of failure?
Because event-chain methods work backwards until they find an anomalous event, and in a system accident there is no anomalous event. Every component did what it was supposed to do. The investigation therefore runs out of anomalies and settles on the nearest human deviation from written procedure, which always exists because real work always deviates from written procedure.
Why is operator error blamed for most accidents?
Leveson's explanation is that after a loss it is always easy to find someone who followed established practice rather than specified practice, because established practice always differs from the written rule. That is why "operator 'error' is found to be the cause of 70-80% of accidents." The figure describes where investigations stop, not where losses originate.
Where do these failures actually happen in a research organisation?
At the boundaries between teams that share an object: the repository written by researchers and read by product managers, the segmentation owned by marketing and used by product, the tracker fielded by an agency and interpreted in-house. Leplat's steel-industry study found that 67 percent of damaging incidents happened in co-activity areas that made up only a small fraction of the total activity space.
What is the practical alternative to blaming a study?
Identify the constraint that was not enforced and the controller who could have enforced it, then fix the feedback that would have told them it was slipping. A constraint with a named owner and an observable signal is a real fix. A reminder is not, and a reminder is what a component-level review produces when there was no defective component.
How do we know whether we have a system accident or just a bad study?
Audit the components first. If a stage of the pipeline has a low first-pass yield, or a batch has a defect rate above your inspection break-even, you have an ordinary component failure and you should fix it. You are in the system-accident class when every study passes its own review, every finding is individually true, and the decision was still wrong.
Related Resources
- Structured Questions Guide - stable question IDs that make cross-study conjunctions auditable
- Research Pipeline Yield - the component-level audit to run before reaching for this framework
- Research Quality Inspection and the All-or-None Rule - why inspection cannot reach interaction failures
- Root Cause Analysis Guide - the right tool for the component-failure class, and its structural limit here
- The Roadmap Evidence Audit - a working procedure for the traceability constraint
- Group Decision Research Reconstruction - what to do when no single person made the decision
- Research Refresh Cadence - half-lives and expiry conditions for the stale-constraint pattern
Related Articles
Nobody Made the Decision: How to Research a Purchase No Single Person Chose (2026)
58% of CEOs' retrospective accounts of their own firm's strategy disagreed with their own earlier validated reports. Why 'why did you buy' is unanswerable for group decisions - and what to measure instead.
Research Pipeline Yield: Why Every Stage Passes and the Finding Still Arrives Wrong
Research stages sit in series, so their pass rates multiply rather than average. Seven stages at 95 percent deliver a correct finding 69.8 percent of the time. How to run a yield audit and fund the lowest stage.
You Cannot Spot-Check Your Way to Data Quality: The All-or-None Rule for Research QA
A ten-item spot check accepts a 5 percent defective batch 59.9 percent of the time. Deming's all-or-none rule says inspect nothing or inspect everything, and sampling is optimal essentially never.
How Long Is User Research Valid? Insight Decay and When to Re-Run a Study
Research does not expire on a fixed schedule — different finding types decay at wildly different rates. A half-life table by insight class, the five decay triggers, and a refresh protocol that keeps your repository honest.
The Roadmap Evidence Audit: Trace Every Claim on Your Roadmap Back to a Source
Take your next ten roadmap items and ask five questions of each. Most teams find their highest-cost commitment is backed by nothing but confident retelling. The audit takes an hour and needs no new tooling.
Root Cause Analysis for Customer Research: The Complete Guide
A practical guide to root cause analysis (RCA) for product and customer research — the 5 Whys, fishbone diagrams, and Pareto analysis — and how to find the real driver behind churn, complaints, and product issues.