{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-09-22T10:58:31.642Z"},"content":[{"type":"documentation","id":"9e68c10c-c2f9-4bb9-8123-52dd3e7cf213","slug":"swiss-cheese-model-research-quality","title":"The Swiss Cheese Model for Research: Why Bad Findings Pass Every Check (2026)","url":"https://www.koji.so/docs/swiss-cheese-model-research-quality","summary":"Applies James Reason's Swiss cheese model (BMJ 2000) to research quality. Flawed findings reach decisions when gaps in several checks line up. Layers only multiply protection when independent: four 80%-effective checks let 0.16% through if independent but 20% if all share the same brief (125x worse). Covers active failures vs latent conditions, a slice-by-slice map of a research study, the OSC 2015 replication result (97% vs 36%), Perneger 2005, and how Koji closes correlated holes.","content":"**Short answer:** A wrong research finding almost never gets through because one check failed. It gets through because several partial checks each had a gap, and on that study the gaps happened to line up. That is the Swiss cheese model, from the psychologist James Reason, and it changes what you fix. Adding one more review rarely helps. You get more from making your existing checks independent of each other, and from finding the *latent conditions* (deadlines, templates, incentives) that open the same gap in every layer at once.\n\nReason set out the model for medicine in a 2000 *BMJ* paper, \"Human error: models and management\". Research teams can take it over almost word for word. A study has layers of defence: the brief, the screener, the interview guide, the moderator, data quality checks, coding, synthesis, peer review, and the stakeholder who reads the deck. Each one catches some problems and misses others. This guide shows how to map those layers, why four checks can protect you no better than one, and where AI-native research tools like Koji close gaps that manual processes leave open.\n\n## What the Swiss cheese model says\n\nReason's picture is a stack of cheese slices. Each slice is a defence, and each has holes. Most of the time a hole in one slice is covered by solid cheese in the next, so nothing gets through. An accident happens only when, in Reason's words, *the holes in many layers momentarily line up to permit a trajectory of accident opportunity*.\n\nTwo details matter more than the cheese itself.\n\n### The holes move\n\nReason is explicit that the defences are *like slices of Swiss cheese, having many holes*, and that unlike in real cheese the holes are *continually opening, shutting, and shifting their location*. The screener that worked last quarter develops a hole when the panel vendor changes its sourcing. Peer review develops a hole the week before a board meeting, when everyone is skimming. A check you tested once is not a check you can count on forever.\n\n### Holes come from two different places\n\nReason separates **active failures** from **latent conditions**. Active failures are *the unsafe acts committed by people who are in direct contact with the patient or system*: slips, lapses, fumbles, mistakes, and procedural violations. Latent conditions are *the inevitable \"resident pathogens\" within the system*: decisions made earlier, often by people far from the work, that sit quietly until they combine with an active failure.\n\nIn research, an active failure is a moderator asking a leading question in interview 7. The latent condition is the reused guide template that already contained that question, or the two-week deadline that meant nobody piloted the guide, or a goal that rewards the team for \"validating\" the roadmap. Active failures are what you see in the transcript. Latent conditions explain why the same failure turns up again next quarter with a different moderator.\n\n## Why \"just add another check\" usually fails\n\nWhen a bad finding reaches a decision, the usual response is to add a layer: a second reviewer, a sign-off step, a new checklist. The model shows why that often buys less than it seems to.\n\nThe arithmetic of layers only works if they are **independent**. Suppose each of four checks catches 80% of a given problem and misses 20%.\n\n| Assumption about the four layers | Share of problems that get through | Roughly |\n|---|---|---|\n| Fully independent (different people, different inputs, different methods) | 0.2 x 0.2 x 0.2 x 0.2 = 0.16% | 1 in 625 |\n| Two independent pairs (each pair shares its inputs) | 0.2 x 0.2 = 4% | 1 in 25 |\n| All four built from the same brief and the same assumption | 20% | 1 in 5 |\n\nThe number of layers is not the number of defences. Four reviewers who all read the same framing, trust the same screener, and use the same codebook are close to one reviewer. Their holes are the same hole. In the table, going from fully independent to fully correlated makes the process **125 times** leakier (20% / 0.16%) with no change in how many people signed off.\n\nResearch makes this correlation hard to avoid, because every layer inherits the brief. If the research question assumes the wrong problem, the screener recruits for it, the guide asks about it, the coder looks for it, and the reviewer checks the deck against it. Every slice passes the finding, correctly by its own standard.\n\n## The evidence that stacked checks can all miss together\n\nThe best-known example is the Open Science Collaboration's 2015 replication project in *Science*. The team repeated 100 studies from three leading psychology journals. **97% of the original studies had reported statistically significant results; only 36% of the replications did**, and replication effect sizes were about half the originals. Every one of those originals had been through design, analysis, significance testing, peer review and editorial review. Those layers were real. They also shared holes: small samples, flexible analysis and a strong preference for novel results all sat upstream of every check.\n\nA second finding matters for anyone who wants to use the model itself. In 2005 Thomas Perneger surveyed 85 quality and safety professionals who said they were very or quite familiar with the Swiss cheese model (*BMC Health Services Research* 5:71). On average they gave **15.3 \"correct\" answers out of 23 (66.5%)** about what the slices, holes and arrow represent, and interpretations *varied considerably*. If safety professionals disagree about what a hole is, a research team using the metaphor loosely will disagree too. Name your slices explicitly, which the next section does.\n\n## Mapping the slices in a research study\n\nHere is a working map for a typical discovery study. The column that matters is the third: what each layer *structurally cannot* catch, however carefully it is done.\n\n| Layer | What it catches | Its characteristic hole | Latent condition that widens the hole |\n|---|---|---|---|\n| Research brief | Wrong question, missing decision | Cannot catch a wrong assumption it was written on | Brief written by the person who owns the roadmap |\n| Screener | Wrong participants | Cannot catch participants who answer screeners strategically | Incentive large enough to attract professional respondents |\n| Interview guide | Leading or missing questions | Cannot catch problems in how questions are delivered | Template reused across studies without piloting |\n| Moderation | Vague answers, missed follow-ups | Cannot catch its own drift over 20 sessions | Back-to-back sessions, one moderator, no calibration |\n| Data quality checks | Speeders, duplicates, junk | Cannot catch fluent, plausible fabrication | Checks tuned for surveys, applied to interviews |\n| Coding and synthesis | Pattern across sessions | Cannot catch themes the codebook has no code for | Codebook drafted before the first interview |\n| Peer review | Overclaiming, weak evidence | Cannot catch problems invisible in the deck | Reviewer sees the summary, never the transcripts |\n| Stakeholder reading | Relevance to the decision | Cannot catch anything it wants to be true | Finding confirms a decision already announced |\n\nTwo things follow from the map. First, the holes are different shapes, which is good: the moderator's hole (drift) is covered by the coder, who reads across sessions. Second, some holes run through every layer. A wrong assumption in the brief goes straight down the stack, and so does a reviewer who only ever sees the summary. Those are the gaps to close first.\n\n## How to apply the model to your research practice\n\n### 1. Write down the slices you actually have\n\nList every check a finding passes before it influences a decision. Most teams find they have fewer than they thought, and that several \"checks\" are the same person looking at the same artifact twice.\n\n### 2. For each slice, state its hole\n\nAsk \"what could get past this layer even if it is done perfectly?\" A screener done perfectly still admits someone who lies well. A peer review done perfectly still misses a quote taken out of context if the reviewer never sees the transcript. Writing the hole down stops people from treating a layer as broader than it is.\n\n### 3. Test for independence\n\nFor every pair of layers, ask whether they share inputs, people, or assumptions. If the same researcher wrote the guide and codes the transcripts, those two slices are correlated. The cheapest fix is often to change one layer's *input*, not to add a layer: give the reviewer three raw transcripts as well as the deck, or have a second person code 20% of sessions blind.\n\n### 4. Hunt latent conditions, not culprits\n\nReason is blunt about blame: *we cannot change the human condition, but we can change the conditions under which humans work*. He reports that *in aviation maintenance... some 90% of quality lapses were judged as blameless*. When a finding turns out to be wrong, tracing it back to the moderator who asked a leading question usually stops one step too early. Ask what let that question into the guide, and whether the same condition is sitting in the next study.\n\n### 5. Trace one failure all the way through\n\nTake one finding that turned out to be wrong and walk it down the stack slice by slice. At each layer, note why it passed. Doing this once teaches a team more about its real defences than any checklist, because it shows where holes lined up in practice rather than in theory.\n\n## How Koji helps close the holes\n\nReason notes that *high reliability organisations... recognise that human variability is a force to harness in averting errors*. The practical question for a research team is which holes should be closed by design so that people's attention goes to the ones only judgment can close. Koji is built to close several of the correlated holes that manual research leaves open.\n\n- **Consistent delivery across every session.** Koji's AI-moderated interviews (text or voice) ask the same core questions the same way to every participant, which closes the moderation-drift hole that lets session 18 quietly become a different study from session 2.\n- **Structured questions that make gaps visible.** Koji supports six structured question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - alongside open conversation. When a key question is structured, a missing or odd answer shows up as a gap in the data instead of disappearing into prose. See the [structured questions guide](/docs/structured-questions-guide).\n- **Evidence trails from finding back to quote.** Koji's automatic thematic analysis links every theme to the participant quotes behind it, so a reviewer can check the transcript as well as the deck. That turns peer review from a correlated slice into an independent one.\n- **Per-interview quality scoring.** Each interview gets a 1-5 quality score, so thin or off-topic sessions are flagged before they are averaged into a theme.\n- **Real-time reporting.** Findings build up as interviews complete, so a hole in the screener or guide shows up at interview 5 rather than after interview 50.\n- **Methodology built into the brief.** Koji's research brief supports frameworks such as Mom Test, JTBD and discovery, with customizable AI consultants that challenge the framing. That puts a check on the one layer every other layer inherits.\n\nA traditional study can take a researcher days of manual review per round. With Koji, consistency checks, quality flags and evidence links are ready within minutes of the last interview. The researcher's time goes to the holes that need a person: whether the question was the right one, and whether the finding is being read honestly.\n\n## Common mistakes when using the Swiss cheese model\n\n- **Counting reviewers as layers.** Three people reading the same deck is one slice with three signatures.\n- **Stopping at the active failure.** Retraining the moderator fixes one hole in one slice. The template that contained the question will produce the same failure with the next moderator.\n- **Treating a slice as fixed.** Holes move. A screener validated a year ago against a different panel is not validated today.\n- **Adding a layer instead of decorrelating one.** A new sign-off step that reads the same summary as the old one adds work without adding protection.\n- **Using the metaphor without defining it.** As Perneger found, even experts read the slices and holes differently. Write down what each slice is and what its hole is.\n- **Relying on tools alone.** Koji closes the delivery, consistency and traceability holes; it cannot tell you whether the brief asked the right question. That slice still needs a human who did not write the brief.\n\n## Frequently asked questions\n\n### What is the Swiss cheese model in simple terms?\n\nIt is a way of thinking about failure in which every safeguard is a slice of cheese with holes. A failure gets through only when the holes in several slices line up. James Reason popularised it, and his 2000 *BMJ* paper \"Human error: models and management\" is the standard reference.\n\n### What is the difference between active failures and latent conditions?\n\nActive failures are errors made at the sharp end, such as a moderator asking a leading question. Latent conditions are problems built into the system earlier, such as a reused template, an unrealistic deadline or an incentive to confirm a decision. Latent conditions can stay hidden for a long time and produce the same active failure again and again.\n\n### How does the Swiss cheese model apply to user research?\n\nEach stage of a study - brief, screener, guide, moderation, data checks, coding, synthesis, review - is a slice. A misleading finding reaches a decision when each stage has a gap in the same place. The model tells you to map those stages, state what each cannot catch, and make them as independent of each other as possible.\n\n### Why doesn't adding more review steps make research more reliable?\n\nBecause extra steps only help if they are independent. If every reviewer works from the same brief and summary, they share the same blind spots, and four correlated checks can protect you no better than one. Changing what a reviewer looks at, for example giving them raw transcripts, often helps more than adding another reviewer.\n\n### Is the Swiss cheese model still considered valid?\n\nIt is still widely used, but it has known limits. Perneger's 2005 survey found that professionals familiar with the model interpreted its parts differently, and later safety researchers have argued it can oversimplify how failures emerge in complex systems. It works best as a shared vocabulary for mapping defences, with each slice and hole defined explicitly.\n\n### How does Koji reduce the chance of holes lining up?\n\nKoji closes several holes by design: consistent AI-moderated delivery, structured questions that expose missing answers, per-interview quality scores, and thematic analysis that links every theme back to quotes. That leaves researchers free to focus on the layers that need judgment, such as whether the research question was right.\n\n## Related Resources\n\n- [Structured Questions in AI Interviews](/docs/structured-questions-guide) - the six question types, and how structure makes a missing answer visible\n- [Research Peer Review as a QA Gate](/docs/research-peer-review-qa-gate) - designing one slice of the stack well\n- [Research Quality Inspection Sampling](/docs/research-quality-inspection-sampling) - why inspection alone cannot carry the quality burden\n- [The Ironies of Automation in Research Analysis](/docs/ironies-of-automation-research-analysis) - why a human reviewer is a weaker slice than it looks\n- [User Research Mistakes](/docs/user-research-mistakes) - the active failures most often found at the sharp end\n- [Triangulation in Research](/docs/triangulation-in-research-guide) - getting genuinely independent evidence rather than correlated checks","category":"Research Methods","lastModified":"2026-09-22T03:19:29.799068+00:00","metaTitle":"Swiss Cheese Model for Research Quality (2026 Guide)","metaDescription":"Why flawed findings pass every research check: the Swiss cheese model, latent conditions, and why four correlated reviews protect no better than one.","keywords":["swiss cheese model","james reason","latent conditions","active failures","research quality","layered defenses","research review process"],"aiSummary":"Applies James Reason's Swiss cheese model (BMJ 2000) to research quality. Flawed findings reach decisions when gaps in several checks line up. Layers only multiply protection when independent: four 80%-effective checks let 0.16% through if independent but 20% if all share the same brief (125x worse). Covers active failures vs latent conditions, a slice-by-slice map of a research study, the OSC 2015 replication result (97% vs 36%), Perneger 2005, and how Koji closes correlated holes.","aiPrerequisites":["Basic familiarity with running a research study","Understanding of peer review or QA steps in research"],"aiLearningOutcomes":["Explain the Swiss cheese model and the difference between active failures and latent conditions","Map the defensive layers in a research study and state each layer's hole","Test research checks for independence rather than counting them","Trace a wrong finding back to its latent conditions"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}