The Swiss Cheese Model for Research: Why Bad Findings Pass Every Check (2026)
How James Reason's Swiss cheese model explains why flawed research findings get past every review, and how to map, decorrelate and strengthen your research defences.
Short answer: A wrong research finding almost never gets through because one check failed. It gets through because several partial checks each had a gap, and on that study the gaps happened to line up. That is the Swiss cheese model, from the psychologist James Reason, and it changes what you fix. Adding one more review rarely helps. You get more from making your existing checks independent of each other, and from finding the latent conditions (deadlines, templates, incentives) that open the same gap in every layer at once.
Reason set out the model for medicine in a 2000 BMJ paper, "Human error: models and management". Research teams can take it over almost word for word. A study has layers of defence: the brief, the screener, the interview guide, the moderator, data quality checks, coding, synthesis, peer review, and the stakeholder who reads the deck. Each one catches some problems and misses others. This guide shows how to map those layers, why four checks can protect you no better than one, and where AI-native research tools like Koji close gaps that manual processes leave open.
What the Swiss cheese model says
Reason's picture is a stack of cheese slices. Each slice is a defence, and each has holes. Most of the time a hole in one slice is covered by solid cheese in the next, so nothing gets through. An accident happens only when, in Reason's words, the holes in many layers momentarily line up to permit a trajectory of accident opportunity.
Two details matter more than the cheese itself.
The holes move
Reason is explicit that the defences are like slices of Swiss cheese, having many holes, and that unlike in real cheese the holes are continually opening, shutting, and shifting their location. The screener that worked last quarter develops a hole when the panel vendor changes its sourcing. Peer review develops a hole the week before a board meeting, when everyone is skimming. A check you tested once is not a check you can count on forever.
Holes come from two different places
Reason separates active failures from latent conditions. Active failures are the unsafe acts committed by people who are in direct contact with the patient or system: slips, lapses, fumbles, mistakes, and procedural violations. Latent conditions are the inevitable "resident pathogens" within the system: decisions made earlier, often by people far from the work, that sit quietly until they combine with an active failure.
In research, an active failure is a moderator asking a leading question in interview 7. The latent condition is the reused guide template that already contained that question, or the two-week deadline that meant nobody piloted the guide, or a goal that rewards the team for "validating" the roadmap. Active failures are what you see in the transcript. Latent conditions explain why the same failure turns up again next quarter with a different moderator.
Why "just add another check" usually fails
When a bad finding reaches a decision, the usual response is to add a layer: a second reviewer, a sign-off step, a new checklist. The model shows why that often buys less than it seems to.
The arithmetic of layers only works if they are independent. Suppose each of four checks catches 80% of a given problem and misses 20%.
| Assumption about the four layers | Share of problems that get through | Roughly |
|---|---|---|
| Fully independent (different people, different inputs, different methods) | 0.2 x 0.2 x 0.2 x 0.2 = 0.16% | 1 in 625 |
| Two independent pairs (each pair shares its inputs) | 0.2 x 0.2 = 4% | 1 in 25 |
| All four built from the same brief and the same assumption | 20% | 1 in 5 |
The number of layers is not the number of defences. Four reviewers who all read the same framing, trust the same screener, and use the same codebook are close to one reviewer. Their holes are the same hole. In the table, going from fully independent to fully correlated makes the process 125 times leakier (20% / 0.16%) with no change in how many people signed off.
Research makes this correlation hard to avoid, because every layer inherits the brief. If the research question assumes the wrong problem, the screener recruits for it, the guide asks about it, the coder looks for it, and the reviewer checks the deck against it. Every slice passes the finding, correctly by its own standard.
The evidence that stacked checks can all miss together
The best-known example is the Open Science Collaboration's 2015 replication project in Science. The team repeated 100 studies from three leading psychology journals. 97% of the original studies had reported statistically significant results; only 36% of the replications did, and replication effect sizes were about half the originals. Every one of those originals had been through design, analysis, significance testing, peer review and editorial review. Those layers were real. They also shared holes: small samples, flexible analysis and a strong preference for novel results all sat upstream of every check.
A second finding matters for anyone who wants to use the model itself. In 2005 Thomas Perneger surveyed 85 quality and safety professionals who said they were very or quite familiar with the Swiss cheese model (BMC Health Services Research 5:71). On average they gave 15.3 "correct" answers out of 23 (66.5%) about what the slices, holes and arrow represent, and interpretations varied considerably. If safety professionals disagree about what a hole is, a research team using the metaphor loosely will disagree too. Name your slices explicitly, which the next section does.
Mapping the slices in a research study
Here is a working map for a typical discovery study. The column that matters is the third: what each layer structurally cannot catch, however carefully it is done.
| Layer | What it catches | Its characteristic hole | Latent condition that widens the hole |
|---|---|---|---|
| Research brief | Wrong question, missing decision | Cannot catch a wrong assumption it was written on | Brief written by the person who owns the roadmap |
| Screener | Wrong participants | Cannot catch participants who answer screeners strategically | Incentive large enough to attract professional respondents |
| Interview guide | Leading or missing questions | Cannot catch problems in how questions are delivered | Template reused across studies without piloting |
| Moderation | Vague answers, missed follow-ups | Cannot catch its own drift over 20 sessions | Back-to-back sessions, one moderator, no calibration |
| Data quality checks | Speeders, duplicates, junk | Cannot catch fluent, plausible fabrication | Checks tuned for surveys, applied to interviews |
| Coding and synthesis | Pattern across sessions | Cannot catch themes the codebook has no code for | Codebook drafted before the first interview |
| Peer review | Overclaiming, weak evidence | Cannot catch problems invisible in the deck | Reviewer sees the summary, never the transcripts |
| Stakeholder reading | Relevance to the decision | Cannot catch anything it wants to be true | Finding confirms a decision already announced |
Two things follow from the map. First, the holes are different shapes, which is good: the moderator's hole (drift) is covered by the coder, who reads across sessions. Second, some holes run through every layer. A wrong assumption in the brief goes straight down the stack, and so does a reviewer who only ever sees the summary. Those are the gaps to close first.
How to apply the model to your research practice
1. Write down the slices you actually have
List every check a finding passes before it influences a decision. Most teams find they have fewer than they thought, and that several "checks" are the same person looking at the same artifact twice.
2. For each slice, state its hole
Ask "what could get past this layer even if it is done perfectly?" A screener done perfectly still admits someone who lies well. A peer review done perfectly still misses a quote taken out of context if the reviewer never sees the transcript. Writing the hole down stops people from treating a layer as broader than it is.
3. Test for independence
For every pair of layers, ask whether they share inputs, people, or assumptions. If the same researcher wrote the guide and codes the transcripts, those two slices are correlated. The cheapest fix is often to change one layer's input, not to add a layer: give the reviewer three raw transcripts as well as the deck, or have a second person code 20% of sessions blind.
4. Hunt latent conditions, not culprits
Reason is blunt about blame: we cannot change the human condition, but we can change the conditions under which humans work. He reports that in aviation maintenance... some 90% of quality lapses were judged as blameless. When a finding turns out to be wrong, tracing it back to the moderator who asked a leading question usually stops one step too early. Ask what let that question into the guide, and whether the same condition is sitting in the next study.
5. Trace one failure all the way through
Take one finding that turned out to be wrong and walk it down the stack slice by slice. At each layer, note why it passed. Doing this once teaches a team more about its real defences than any checklist, because it shows where holes lined up in practice rather than in theory.
How Koji helps close the holes
Reason notes that high reliability organisations... recognise that human variability is a force to harness in averting errors. The practical question for a research team is which holes should be closed by design so that people's attention goes to the ones only judgment can close. Koji is built to close several of the correlated holes that manual research leaves open.
- Consistent delivery across every session. Koji's AI-moderated interviews (text or voice) ask the same core questions the same way to every participant, which closes the moderation-drift hole that lets session 18 quietly become a different study from session 2.
- Structured questions that make gaps visible. Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - alongside open conversation. When a key question is structured, a missing or odd answer shows up as a gap in the data instead of disappearing into prose. See the structured questions guide.
- Evidence trails from finding back to quote. Koji's automatic thematic analysis links every theme to the participant quotes behind it, so a reviewer can check the transcript as well as the deck. That turns peer review from a correlated slice into an independent one.
- Per-interview quality scoring. Each interview gets a 1-5 quality score, so thin or off-topic sessions are flagged before they are averaged into a theme.
- Real-time reporting. Findings build up as interviews complete, so a hole in the screener or guide shows up at interview 5 rather than after interview 50.
- Methodology built into the brief. Koji's research brief supports frameworks such as Mom Test, JTBD and discovery, with customizable AI consultants that challenge the framing. That puts a check on the one layer every other layer inherits.
A traditional study can take a researcher days of manual review per round. With Koji, consistency checks, quality flags and evidence links are ready within minutes of the last interview. The researcher's time goes to the holes that need a person: whether the question was the right one, and whether the finding is being read honestly.
Common mistakes when using the Swiss cheese model
- Counting reviewers as layers. Three people reading the same deck is one slice with three signatures.
- Stopping at the active failure. Retraining the moderator fixes one hole in one slice. The template that contained the question will produce the same failure with the next moderator.
- Treating a slice as fixed. Holes move. A screener validated a year ago against a different panel is not validated today.
- Adding a layer instead of decorrelating one. A new sign-off step that reads the same summary as the old one adds work without adding protection.
- Using the metaphor without defining it. As Perneger found, even experts read the slices and holes differently. Write down what each slice is and what its hole is.
- Relying on tools alone. Koji closes the delivery, consistency and traceability holes; it cannot tell you whether the brief asked the right question. That slice still needs a human who did not write the brief.
Frequently asked questions
What is the Swiss cheese model in simple terms?
It is a way of thinking about failure in which every safeguard is a slice of cheese with holes. A failure gets through only when the holes in several slices line up. James Reason popularised it, and his 2000 BMJ paper "Human error: models and management" is the standard reference.
What is the difference between active failures and latent conditions?
Active failures are errors made at the sharp end, such as a moderator asking a leading question. Latent conditions are problems built into the system earlier, such as a reused template, an unrealistic deadline or an incentive to confirm a decision. Latent conditions can stay hidden for a long time and produce the same active failure again and again.
How does the Swiss cheese model apply to user research?
Each stage of a study - brief, screener, guide, moderation, data checks, coding, synthesis, review - is a slice. A misleading finding reaches a decision when each stage has a gap in the same place. The model tells you to map those stages, state what each cannot catch, and make them as independent of each other as possible.
Why doesn't adding more review steps make research more reliable?
Because extra steps only help if they are independent. If every reviewer works from the same brief and summary, they share the same blind spots, and four correlated checks can protect you no better than one. Changing what a reviewer looks at, for example giving them raw transcripts, often helps more than adding another reviewer.
Is the Swiss cheese model still considered valid?
It is still widely used, but it has known limits. Perneger's 2005 survey found that professionals familiar with the model interpreted its parts differently, and later safety researchers have argued it can oversimplify how failures emerge in complex systems. It works best as a shared vocabulary for mapping defences, with each slice and hole defined explicitly.
How does Koji reduce the chance of holes lining up?
Koji closes several holes by design: consistent AI-moderated delivery, structured questions that expose missing answers, per-interview quality scores, and thematic analysis that links every theme back to quotes. That leaves researchers free to focus on the layers that need judgment, such as whether the research question was right.
Related Resources
- Structured Questions in AI Interviews - the six question types, and how structure makes a missing answer visible
- Research Peer Review as a QA Gate - designing one slice of the stack well
- Research Quality Inspection Sampling - why inspection alone cannot carry the quality burden
- The Ironies of Automation in Research Analysis - why a human reviewer is a weaker slice than it looks
- User Research Mistakes - the active failures most often found at the sharp end
- Triangulation in Research - getting genuinely independent evidence rather than correlated checks
Related Articles
The Ironies of Automation: Why a Human Reviewer Cannot Catch Your AI's Analysis Errors (2026)
Adding a human to spot-check AI coding is the reflex fix. Bainbridge showed in 1983 why it backfires, and the arithmetic is worse than teams expect.
Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies
Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.
You Cannot Spot-Check Your Way to Data Quality: The All-or-None Rule for Research QA
A ten-item spot check accepts a 5 percent defective batch 59.9 percent of the time. Deming's all-or-none rule says inspect nothing or inspect everything, and sampling is optimal essentially never.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.
Triangulation in Research: Combining Methods for Stronger, More Credible Insights (2026)
Triangulation is the practice of using multiple data sources, methods, researchers, or theories to validate a finding. Learn Denzin's four types, when to use each, and how AI-native research platforms make multi-method studies practical instead of aspirational.
User Research Mistakes: 14 Pitfalls That Sabotage Your Insights (2026)
The most common user research mistakes that lead to misleading insights — and how to avoid each one with better methodology and AI-powered interviews.