Interim Analysis and Stopping Rules: How to Stop a Study Early Without Faking the Result (2026)
Clinical trials solved the problem of looking at data before a study ends. Group sequential designs, alpha spending, and futility boundaries let you stop early without inflating false positives - and the evidence shows what happens when you stop early without them.
Answer first: stopping a study early is legitimate if, and only if, you fixed the rules for stopping before you saw any data. Clinical trials have operated this way for nearly fifty years and have the apparatus to prove it: group sequential designs, alpha-spending functions, pre-specified futility boundaries, and an independent committee that holds the interim results so the research team cannot see them. Product research has almost none of this. That is why "the lift was obvious by day three so we called it" and "we kept fielding until the numbers firmed up" are the same act described in different tones of voice.
This guide ports that apparatus into product and UX research at a scale a product team can actually operate. It is deliberately not the standard advice.
The standard advice is correct, and it is also too expensive to follow
The received wisdom in research methods is simple: decide your sample size in advance, do not look until you get there, and do not stop when the result turns pretty. That rule is right about the danger. Checking a result repeatedly and stopping the moment it crosses significance is one of the best-documented ways to manufacture a false positive, and our guide to p-hacking and researcher degrees of freedom covers exactly why: every additional look is another chance for noise to cross the line, and the nominal 5 percent error rate quietly stops meaning 5 percent.
But look at what the advice actually asks of a product team. It asks you to keep running a study you already suspect is pointless. It asks you to keep recruiting participants for an effect that is clearly not there. It asks you to spend two more weeks on a question that was answered in week one. In a clinical trial, refusing to look has a body count - you may be withholding a treatment that works, or continuing to administer one that harms. In product research the cost is smaller but real: burned participant goodwill, a research team stuck on a dead question, and a decision made late.
Medicine did not resolve this tension by telling investigators never to look. It resolved it by building a formal accounting system for looking. That system is the thing product research never imported.
The evidence: trials stopped early systematically overstate the effect
Two systematic reviews in JAMA define the problem, and both are worth knowing in detail because their findings are more subtle than the headline.
Montori and colleagues (JAMA 294(17):2203-2209, 2005) catalogued 143 randomised trials reported as stopped early for benefit. The majority - 92 of them - were published in five high-impact medical journals. These trials had recruited on average 63 percent of their planned sample (SD 25 percent) and stopped after a median of 13 months of follow-up, a median of one interim analysis, and when a median of 66 patients had experienced the event that triggered termination (interquartile range 23 to 195). The median risk ratio among them was 0.53 (IQR 0.28 to 0.66) - an apparent halving of risk.
The reporting was worse than the statistics. 135 of the 143 trials (94 percent) failed to report at least one of the planned sample size, which interim analysis prompted the stop, whether a stopping rule informed the decision at all, or an analysis adjusted for the interim monitoring. Trials with fewer events produced larger treatment effects, with an odds ratio of 28 (95 percent CI 11 to 73). The authors concluded that clinicians should view results of such trials with skepticism.
Bassler and colleagues (JAMA 303(12):1180-1187, 2010) then did the decisive comparison. They matched 91 truncated trials, asking 63 distinct questions, against 424 trials asking the same questions that were not stopped early. The pooled ratio of relative risks was 0.71 (95 percent CI 0.65 to 0.77) - trials stopped early reported effects roughly 29 percent larger than the trials that ran to completion. The gap was widest when the truncated trial had fewer than 500 events.
And then the finding that matters most for anyone about to write a stopping rule into a research plan: in 39 of the 63 questions (62 percent), the pooled evidence from the trials that were not stopped early failed to demonstrate a significant benefit at all. The early stop did not just exaggerate the size of the effect. In most cases the effect was not reliably there.
| What the reviews measured | Montori 2005 (n=143) | Bassler 2010 (n=91 vs 424) |
|---|---|---|
| Fraction of planned sample recruited | 63 percent (SD 25) | not reported |
| Typical number of interim looks before stopping | median 1 | not reported |
| Events at the stop | median 66 (IQR 23-195) | overstatement worst under 500 |
| Effect size reported | median RR 0.53 | RRR 0.71 vs non-truncated |
| Reporting failures | 94 percent omitted key details | not the focus |
| Effect held up in later evidence | not assessed | no significant benefit in 62 percent of questions |
The uncomfortable part: having a stopping rule did not rescue the estimate
This is the finding that separates a serious treatment of this topic from a superficial one, and it is the reason this guide does not simply tell you to write a stopping rule and move on.
Bassler and colleagues tested whether the overstatement was smaller in trials that had a formal statistical stopping rule. It was not. The difference was independent of the presence of a statistical stopping rule and independent of methodological quality as measured by allocation concealment and blinding. What predicted the overstatement was the number of events - that is, how little information the trial had accumulated when it stopped.
The correct conclusion is not that stopping rules are useless. It is that a stopping rule controls the false positive rate - the probability of crossing a boundary by chance - and does nothing about effect size inflation, which is a separate phenomenon. If you stop the first time a noisy estimate wanders above a threshold, you have selected a high draw from the sampling distribution. The boundary made that selection legitimate in one narrow sense and left the magnitude biased upward regardless.
Practically: a study stopped early tells you the direction is probably real; it does not tell you how big the effect is, and you should not plan on the observed magnitude. If you stopped early on a 30 percent lift, budget as though it were meaningfully smaller. This is the same logic that makes minimum detectable effect worth calculating in advance, and the same logic behind publication bias: the results that reach you have already been filtered by a process that favours large numbers.
The machinery: three families of boundary
The statistical solution is to spend your error budget across looks instead of pretending you only looked once. Three approaches dominate, and the differences between them are practical rather than academic.
| Approach | How it allocates error | Behaviour | When it fits product research |
|---|---|---|---|
| Pocock (1977) | Same nominal threshold at every look | Easy to stop early; pays for it with a stricter final threshold | Studies where an early answer is genuinely much more valuable than a precise one |
| O'Brien-Fleming (1979) | Very strict early, near-normal at the end | Early stops require an overwhelming result; the final analysis is barely penalised | The sensible default for most product studies |
| Lan-DeMets alpha spending (1983) | A function that spends error continuously | The number and timing of looks does not have to be fixed in advance | Continuous or always-on research where you cannot schedule looks |
Pocock proposed the constant-boundary group sequential design in Biometrika in 1977. O'Brien and Fleming published their multiple testing procedure in Biometrics 35:549-556 in 1979; it is conservative early and permissive late, which is why it became the default in large trials. Lan and DeMets then generalised the idea in 1983 into an alpha-spending function, removing the requirement that you commit in advance to exactly how many interim analyses you will run - you commit instead to the rate at which you consume your error budget.
For a product team, the intuition transfers even if the arithmetic is delegated to software. O'Brien-Fleming is the right mental model: to stop a study in its first week, the result should be so lopsided that no reasonable person would want more data. To stop it at the scheduled end, the ordinary threshold applies. That single asymmetry prevents most of the damage, because it removes the incentive to call a marginal early result.
Futility: the stopping decision nobody in product ever makes
Every product team has stopped a study early because the result looked good. Almost none have stopped one because the result looked like nothing.
Clinical trials formalise this as a futility boundary (sometimes called a non-binding futility analysis): a pre-specified threshold at which the accumulated data are so unpromising that continuing cannot plausibly reach a useful answer. Crossing it means you stop and free the resources. This is the mirror image of the efficacy boundary and it is by far the more valuable of the two for product research, because the base rate of product ideas that do nothing is high. Our guide to quasi-experimental design cites the Microsoft finding that only about one third of tested ideas improve the metric they were designed to improve. If two thirds of your studies are heading nowhere, a rule that ends them at the halfway mark is the single highest-return process change available to you.
Futility stopping is also statistically benign in a way efficacy stopping is not. Stopping early for benefit inflates the effect you report. Stopping early for futility produces a null you were going to get anyway - and if you want to make a positive claim of no meaningful difference rather than merely failing to find one, that is a different test entirely, covered in equivalence testing.
Who is allowed to look
The part of the clinical-trials apparatus that transfers most cleanly is not statistical at all. It is organisational.
A trial that permits interim analysis does not let the investigators see the interim results. An independent data and safety monitoring board sees them, applies the pre-specified rules, and reports back only a recommendation: continue, modify, or stop. The investigators stay blind. The reason is not distrust of individuals; it is that knowing the interim direction contaminates every subsequent judgement call - who gets recruited next, how hard a protocol deviation is scrutinised, how an ambiguous case is coded.
The underlying ethical principle is equipoise: you are only entitled to run the study while genuine uncertainty about the answer exists. When it stops existing, continuing is no longer justified. Product research has a weaker version of the same obligation. Continuing to recruit participants into a study whose answer you already have is a waste of their time and of your relationship with them.
The transferable practice: separate the person who reads the interim cut from the person who runs the study and owns the hypothesis. In a small team this is one colleague from another squad, given the rule in advance and asked to return a one-word recommendation. That is a fifteen-minute imposition and it converts an unaccountable peek into an auditable decision. It is the same separation-of-duties logic behind blind analysis, where the analyst works on data whose group labels are masked until the analysis plan is locked.
Porting this to product research
Quantitative studies
A survey or an experiment can take the machinery almost literally. Pre-specify the primary outcome, the target sample, the number and timing of interim looks, the boundary family, and the futility rule. Write it down before fielding. If you are checking a live dashboard daily without any of that, you are running an uncontrolled sequential design and your stated confidence intervals do not mean what they say - see statistical significance in survey research and the multiple comparisons problem for what that does to your error rate.
Qualitative studies
Qualitative research already has a stopping rule; it is just rarely treated as one. Saturation is a stopping rule with no error budget attached. "We stopped when we stopped hearing new themes" is a decision made by the same person who wants the study to be finished, evaluated against no pre-specified criterion, on data they have already read. Our guide to data saturation covers how to operationalise it properly, and how many interviews are enough covers the sample-size question underneath it.
The upgrade is to specify saturation in advance in countable terms: for example, "we stop when three consecutive interviews produce no code that is not already in the codebook, with a floor of 12 and a ceiling of 25." The floor prevents an early stop on a homogeneous run; the ceiling forces a decision instead of open-ended drift.
| Study type | Pre-specify this | Efficacy stop | Futility stop |
|---|---|---|---|
| Survey with a primary metric | Target n, look schedule, boundary family | Result crosses the early boundary | Confidence interval already excludes the effect you would act on |
| Concept test across variants | Primary variant comparison, look schedule | One variant wins decisively at the interim look | No variant separates from control by the midpoint |
| Qualitative discovery | Codebook, new-code rule, floor and ceiling | Three consecutive interviews add no new code | Ceiling reached; write up what you have |
| Continuous discovery programme | Review cadence, decision the research feeds | The decision it informs has been made | The decision was cancelled or made without it |
That last row is the one continuous programmes miss. A weekly interview cadence with no termination condition is not a study; it is a subscription. Continuous discovery is more useful when each thread inside it has a defined end.
The modern approach: how Koji helps
The reason product teams peek is not indiscipline. It is that traditional research is slow enough that waiting for the full sample is genuinely costly, so the incentive to call it early is enormous. Compressing the timeline removes most of the temptation at source.
Speed collapses the window in which peeking pays. Koji runs AI-moderated interviews in parallel rather than in sequence, so a study that would take a human moderator three weeks of scheduled sessions can field in days. When the full sample arrives before your decision meeting, the early stop stops being attractive.
Pre-specification lives in the research brief, not in someone's memory. Koji structures a study around an explicit brief with a methodology - mom_test, jtbd, discovery, exploratory, or lead_magnet - and a defined question set written before fielding begins. Writing the target sample and the stopping rule into that brief makes them an artifact with a timestamp rather than a claim made afterwards.
Structured questions make an interim look countable rather than impressionistic. Koji supports six question types - open_ended, scale, single_choice, multiple_choice, ranking, and yes_no - and each carries a stable question ID from the interview plan through moderation and analysis into the report. That means an interim cut is a comparison of the same identified items, not a recollection of how the last few conversations felt. See the structured questions guide for how the six types map to reportable outputs.
Consistent moderation removes the collection-side drift that interim looks normally introduce. A human moderator who has seen a promising interim pattern probes harder for it in the next session. An AI moderator runs the same protocol on interview 5 and interview 50, so the data collected after the look is comparable to the data collected before it - which is precisely the property that makes sequential analysis valid.
Real-time reporting makes futility visible early. Because analysis is continuous rather than a task that begins after fieldwork ends, the case for stopping a study that is going nowhere is available at the point where stopping still saves something.
The honest framing: Koji does not compute O'Brien-Fleming boundaries for you. It removes the operational pressure that makes undisciplined peeking attractive, and it makes the pre-specified plan and the interim comparison into durable, inspectable objects. The discipline is still yours.
Frequently asked questions
Is it ever acceptable to stop a study early because the result is already clear?
Yes, provided the definition of "clear" was written down before you saw data, and provided you treat the observed effect size as an overestimate. Bassler and colleagues found trials stopped early reported effects about 29 percent larger than comparable trials that ran to completion, and that gap persisted even in trials that had formal stopping rules. Stop if you must, but plan against a smaller number than the one you observed.
What is the difference between a stopping rule and p-hacking?
Timing and specificity. A stopping rule fixed in advance - "we field until 400 completes or 14 days, whichever comes first" - constrains your behaviour. A decision made after seeing the data, about data you have already seen, does not. Optional stopping is one of the best-documented researcher degrees of freedom, and our guide to p-hacking covers how much it inflates false positives on its own.
What is a futility boundary and why does it matter more than an efficacy boundary?
A futility boundary is a pre-set threshold at which the data are unpromising enough that continuing cannot plausibly produce a useful answer, so you stop. It matters more in product research because most tested ideas do not move the metric they targeted, so most studies are candidates for it - and unlike stopping for benefit, stopping for futility does not bias the effect size you report.
How many interim looks should a product study have?
For most studies, one - at roughly the midpoint - is the right answer. One look captures most of the practical benefit (the chance to stop something futile halfway through) with the smallest cost to your error budget. Montori and colleagues found the trials in their review had a median of exactly one interim analysis before stopping.
Does saturation in qualitative research count as a stopping rule?
It does, and that is the problem: it is usually applied with no pre-specified criterion, by the person who wants the study finished, on data they have already read. Convert it into something countable in advance - a new-code rule with a floor and a ceiling - and it becomes a real stopping rule rather than a post-hoc justification.
Who should look at the interim results?
Ideally not the person who owns the hypothesis. Clinical trials use an independent monitoring board precisely so investigators stay blind to interim direction. In a product team, hand the pre-specified rule to a colleague outside the squad and ask them to return a recommendation rather than a number. It costs about fifteen minutes and turns an unaccountable peek into an auditable decision.
Related Resources
- P-Hacking and Researcher Degrees of Freedom - why unplanned peeking inflates false positives, and what pre-specification fixes
- Statistical Power and Minimum Detectable Effect - deciding in advance what size of effect your study can actually see
- Data Saturation in Qualitative Research - turning "we stopped hearing new things" into a defensible criterion
- Equivalence Testing - how to claim there is no meaningful difference rather than merely failing to find one
- Blind Analysis - separating the analyst from the answer so the interim look cannot steer the analysis
- Structured Questions in AI Interviews - the six question types that make an interim cut countable
Related Articles
Blind Analysis: How to Analyze Research Before You Know the Answer
Blind analysis hides which group is which until your analysis is locked. Borrowed from particle physics, it is the cheapest way to stop your expectations from steering your findings.
Data Saturation in Qualitative Research: How to Know When You Have Enough
Data saturation is the point at which additional interviews stop producing new information. This guide covers the four types of saturation (theoretical, data, code, meaning), how to recognize and document them, the empirical sample sizes from Hennink and Guest, and how AI-moderated interviews let you reach saturation in days instead of months.
How to Prove There Is No Difference: Equivalence Testing for Product Research (2026)
A non-significant result does not mean there is no difference - it usually means your study could not tell. Equivalence testing is the method that lets you actually claim two things are the same, and product teams make expensive no-difference decisions without it every quarter.
The Multiple Comparisons Problem: Why Slicing Data Into Segments Manufactures Findings (2026)
Test 20 segments at the 5 percent threshold and you have a 64 percent chance of finding at least one difference that is not there. Learn how to count the tests you actually ran, when to control the family-wise error rate versus the false discovery rate, and why a correction cannot rescue a bad prior.
P-Hacking and Researcher Degrees of Freedom: How Analytic Flexibility Manufactures Findings (2026)
Four ordinary analytic choices raise the false-positive rate from 5 percent to 61 percent. Learn what researcher degrees of freedom are, why the garden of forking paths catches honest researchers, and how a one-page pre-committed analysis plan fixes it without banning exploration.
Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)
Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.
Statistical Significance in Survey Research: A Plain-English Guide (2026)
A plain-English guide to statistical significance for survey and market researchers: what p-values and confidence levels really mean, how to test differences, the myths to avoid, and when significance matters less than insight.
Structured Questions in AI Interviews
Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.