Back to docs
Research Methods

Futility Analysis: How to Decide a Running Study Will Never Answer Your Question

Most teams only ever ask whether a study can be stopped early because it worked. The more valuable question is whether it can be stopped because it never will. Futility rules, stop-for-harm, and who is allowed to make the call.

Answer first: there are exactly three reasons to stop a study before its planned end - it worked, it will never work, or it is causing harm. Product teams routinely act on the first, almost never plan for the second, and have no process at all for the third. That ordering is backwards. The evidence from clinical trials is that stopping early for benefit produces systematically overstated effects: in a systematic review of 91 truncated trials matched against 424 trials that ran to completion, the pooled ratio of relative risks was 0.71, and in 39 of the 63 clinical questions studied (62%) the full non-truncated evidence base failed to show a significant benefit at all. Stopping early because it will never work, by contrast, costs nothing but the budget you save, and it is the decision nobody schedules.

This guide covers futility rules, the stop-for-harm case that product research has no vocabulary for, and the governance question underneath both: who gets to call it.

The three stopping reasons, and how differently they behave

Clinical trials are monitored on all three continuously. The asymmetry between them is the single most useful thing to import.

Reason to stopWhat it claimsEvidence quality when you act on itHow often product teams plan for it
EfficacyThe effect is real and large enough to act nowSystematically overstated, worst with few eventsConstantly, informally
FutilityThe study will not reach a usable answerRobust - you are declining to claim anythingAlmost never
HarmThe thing being tested is hurting peopleUsually decisiveNo process at all

Stopping for success is the dangerous one

Two systematic reviews establish this and they are worth knowing by number.

Montori and colleagues (JAMA, 2005, 294(17):2203-2209) identified 143 randomized trials stopped early for benefit, 92 of them published in five high-impact medical journals. The proportion of trials in those journals stopped early for benefit rose from 0.5% in 1990-1994 to 1.2% in 2000-2004. On average these trials recruited 63% of their planned sample and stopped after a median of 13 months of follow-up, one interim analysis, and a median of just 66 events. The median risk ratio was 0.53 - an apparent halving of risk. And 135 of the 143 (94%) failed to report at least one of: the planned sample size, the interim analysis after which the trial stopped, whether a stopping rule informed the decision, or an adjusted analysis accounting for truncation. That last item was missing in 129 of 143. Trials with fewer events reported larger effects, with an odds ratio of 28 (95% CI 11 to 73).

Bassler and colleagues (JAMA, 2010, 303(12):1180-1187) then compared 91 truncated trials against 424 matched trials that were not stopped early. The pooled ratio of relative risks was 0.71 (95% CI 0.65 to 0.77) - truncated trials reported effects about 29% larger. Critically, this difference was independent of whether a statistical stopping rule was present and independent of methodological quality. The overstatement was worst in trials with fewer than 500 events. And in 39 of 63 questions (62%), the pooled effect from the non-truncated trials failed to demonstrate significant benefit at all.

The translation to product research is direct. A study stopped at 40% of target because the early numbers looked strong is the research equivalent of a truncated trial, with far fewer observations than 500 events and no adjustment. Whatever effect size it reports should be treated as an upper bound, not an estimate. If you must stop early for success, do it against a pre-declared boundary - see interim analysis and sequential testing for the thresholds - and report the effect as provisional.

Futility: the decision that saves the most money

Futility asks a different question. Not "is the effect real" but "given what we have so far, is there any plausible way the remaining sample changes the conclusion?"

The standard tool is conditional power: the probability that the study reaches its threshold at the planned end, given the data already collected and an assumption about the true effect. If conditional power is very low - the conventional trigger is somewhere in the 10% to 20% range - continuing is spending budget to confirm something you already know.

Futility has a property that makes it unusually safe to act on: you are declining to make a claim, not making one. The overstatement problem that afflicts efficacy stopping does not apply, because there is no effect estimate being published. The worst case is that you abandoned a study that would have squeaked over the line, which is a cost, not an error.

Futility rules that work without conditional power

Most customer research cannot compute conditional power, and does not need to. Futility in practice is usually structural rather than statistical, and these four rules cover the great majority of wasted studies:

  1. Recruitment futility. The study needs 120 responses from a segment that has produced 4 in two weeks. No analysis will fix a sample that does not exist. Set the rule as a rate: "if we are below 40% of target at the halfway date, we stop and redesign recruitment."
  2. Variance futility. The outcome is so noisy that the confidence interval at full sample would still contain both "large improvement" and "no change." This is a power problem discovered late; see statistical power and minimum detectable effect for how to catch it before fielding.
  3. Question futility. Respondents are not answering the question you asked. If the open_ended responses to your key item are consistently about something else, more of them will not help. The fix is a new instrument, not a bigger sample.
  4. Decision futility. The most underrated one. Ask, at the halfway look: if the result comes back at the most favourable plausible value, does anyone change what they were going to do? If not, the study is futile regardless of its statistics, and it was futile before it launched. Our guide to research peer review covers catching this at the pre-launch gate, which is where it belongs.

Stopping for harm: the case product research has no words for

The definitive example is the Cardiac Arrhythmia Suppression Trial (CAST), reported in the New England Journal of Medicine in 1989 (321(6):406-412). The premise was sound: ventricular premature depolarizations after a heart attack predict sudden death, so suppressing them should save lives. Of 2,309 patients recruited to the titration phase, 1,727 (75%) had their arrhythmia successfully suppressed by one of the study drugs and were randomized to active drug or placebo.

The drugs worked. The surrogate endpoint moved exactly as intended. Over an average of 10 months, patients on encainide or flecainide had 33 arrhythmic deaths or cardiac arrests out of 730 (4.5%), against 9 out of 725 on placebo (1.2%) - a relative risk of 3.6 (95% CI 1.7 to 8.5). Total mortality was 56 of 730 (7.7%) versus 22 of 725 (3.0%), a relative risk of 2.5 (95% CI 1.6 to 4.5). That arm of the trial was stopped.

CAST is the canonical demonstration that a surrogate metric can move in the right direction while the outcome you actually care about moves in the wrong one. Every product team that optimises engagement while retention quietly falls, or reduces support tickets by making the help page harder to find, is running a small CAST. The reason clinical trials catch it and product research does not is not superior ethics; it is that somebody is explicitly tasked with watching the harm endpoint, and that person is not the one who wants the treatment to work.

In research specifically, harm has a narrower and very real meaning: a study can hurt its participants. Questions that surface trauma, screeners that expose sensitive status, an incentive structure that pressures people into completing, or an interview that runs three times its advertised length. A harm-stopping rule for a research study is a sentence in the brief naming the signal that halts fielding immediately - a distress report, a complaint pattern, an unexpected drop-off cliff on a sensitive item.

Who is allowed to call it

This is the part that transfers most cleanly and that almost nobody has implemented.

Clinical trials use a Data Monitoring Committee (also called a DSMB or DMC). It is typically a committee of three to nine external experts - clinicians, one or two statisticians, and often an ethicist - who meet once or twice a year, review trial conduct, safety, futility and efficacy, and recommend to the sponsor whether to continue, modify or stop. The sponsor makes the final decision but in practice almost always accepts the recommendation.

The FDA draft guidance Use of Data Monitoring Committees in Clinical Trials (February 2024) states the structural principle plainly: a DMC "is established by the sponsor but should be independent of the sponsor and the trial conduct." And on the specific decision this guide is about: "Changes to the trial design that involve an analysis of results by study group are best performed by a body independent of the sponsor, the investigators, and the subjects."

Note what that sentence rules out. The person who proposed the feature, the person running the study, and the person who will present the results should not be the person deciding, on the basis of interim results, whether to stop. In most product teams that is one person, and it is usually the same person who is watching the dashboard every morning.

A lightweight version that fits a product org

You do not need a nine-person committee for a 60-participant interview study. You need the firewall, not the ceremony.

ElementClinical versionProduct research version
Who monitorsExternal DMC, 3-9 membersOne named reviewer outside the requesting team
What they seeUnblinded comparative interim dataComparative interim results, which nobody else sees
What everyone else seesAggregate, blinded reportsOperational metrics only - pace, completion, segment fill
Decision authorityRecommends to sponsorRecommends to the study owner, in writing
Governing documentDMC charterThree lines in the study brief
Conflict ruleNo ongoing financial relationship with sponsorReviewer does not own the roadmap item being tested

The three lines in the brief are: who may see comparative interim results, what the futility and harm triggers are, and who signs off on stopping. That is the whole intervention. It costs nothing and it removes the single most common way a research programme talks itself into a conclusion.

How Koji supports stopping decisions

Operational monitoring is separable from outcome analysis. Koji reports show recruitment pace, completion status and per-question drop-off alongside findings, so the reviewer watching for recruitment futility does not have to read the substantive results to do their job. That separation is what makes a firewall practical rather than theoretical.

Parallel fielding shortens the window in which futility is expensive. Because AI-moderated voice and text interviews run concurrently rather than one at a time, a study that would have taken three weeks of sequential scheduling completes in days. Futility discovered on day two costs two days. The same futility discovered in week three of manual scheduling costs three weeks and most of the budget.

Structured questions make futility triggers checkable. With all six question types available - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - the primary outcome can be a specific item with a defined distribution, so a futility rule like "if fewer than 15% select this option at halfway, we stop" is something a monitor can evaluate without interpretation. See the structured questions guide. A study whose only outcome is free text has no checkable futility trigger, which is one reason free-text-only studies run long.

Credits follow completed quality conversations. Koji applies a quality gate so that only conversations meeting a quality threshold consume credits, which means a study stopped for futility does not bill for the low-quality tail it was collecting on its way to nowhere.

Compared with traditional survey platforms such as SurveyMonkey, Typeform or Qualtrics - where a fielded survey simply runs until its close date and the only monitoring surface is a response counter - platforms like Koji give a monitor enough structure to act on a stopping rule while the study is still cheap to stop.

Frequently asked questions

What is the difference between futility and just not having enough data yet?

Futility is a forward-looking judgment: given what you have, the remaining planned sample cannot plausibly change the conclusion. Not having enough data yet means the remaining sample can change it and you should keep going. The test is arithmetic in quantitative work (conditional power) and structural in qualitative work (can the recruitment, the instrument or the decision context deliver an answer at all).

Is stopping a study early for good results ever acceptable?

Yes, when the boundary was set in advance, the study crosses it, and the readout reports the effect as an upper bound rather than a point estimate. The evidence is unambiguous that truncated studies overstate effects - by roughly 29% on average in the Bassler analysis - and that the overstatement persists even when a formal stopping rule was in place. Treat an early win as a reason to plan a confirmatory study, not as the confirmation.

Who should sit on a monitoring role in a small company?

Someone who does not own the outcome. A researcher from an adjacent team, a data analyst, or the person who ran the pre-launch peer review. The only hard requirement is that stopping the study early, in either direction, must not affect anything they are accountable for.

How do you set a futility rule for a qualitative study?

Use recruitment and instrument criteria rather than statistical ones. For example: "if we have not completed 12 interviews in the primary segment by day 7, we stop and revisit screening" or "if three consecutive participants cannot answer the core question as written, we stop and rewrite it." Both are checkable, and both catch the two ways qualitative studies actually fail.

What is a harm-stopping rule in customer research?

A named signal that halts fielding immediately, written into the brief before launch. Common ones: any participant reporting distress, a complaint pattern about a specific question, an unexpected drop-off cliff on a sensitive item, or discovery that a screener is exposing information participants did not consent to share. The point of writing it down is that in the moment, the person who sees the signal is usually the person least empowered to stop the study.

Does this apply to always-on or continuous research programmes?

Yes, and futility matters more there, because a continuous programme has no natural end date to force the question. Schedule a standing futility review - quarterly is typical - that asks whether the programme is still changing decisions. Decision futility is the failure mode of continuous discovery, and it is invisible without a scheduled check.

Related Resources

Related Articles

How to Prove There Is No Difference: Equivalence Testing for Product Research (2026)

A non-significant result does not mean there is no difference - it usually means your study could not tell. Equivalence testing is the method that lets you actually claim two things are the same, and product teams make expensive no-difference decisions without it every quarter.

P-Hacking and Researcher Degrees of Freedom: How Analytic Flexibility Manufactures Findings (2026)

Four ordinary analytic choices raise the false-positive rate from 5 percent to 61 percent. Learn what researcher degrees of freedom are, why the garden of forking paths catches honest researchers, and how a one-page pre-committed analysis plan fixes it without banning exploration.

Publication Bias and the File-Drawer Problem in Product Research: Why Your Evidence Base Only Remembers the Studies That Worked (2026)

Publication bias is not an academic curiosity. In product research it is worse, because nobody rejects your null study - you simply never write it up. Learn how big the file drawer is, what it does to your confidence, and how to build a study register that closes it.

Research Peer Review: The Pre-Launch QA Gate That Catches Broken Studies

Most research quality programmes police respondents. Almost none police the study design. A 30-minute structured review before fieldwork catches the errors that no amount of data cleaning can fix afterwards.

Statistical Power and Minimum Detectable Effect: Can Your Survey Detect the Change You Care About? (2026)

Margin of error tells you how precise one number is. Minimum detectable effect tells you how big a change has to be before you can see it — and it is roughly twice as large. Includes MDE tables for proportions, scales and NPS.

5-Point vs 7-Point Likert Scale: How Many Scale Points Should You Use? (2026)

A decision guide for rating-scale length — what the reliability research actually says about 5 vs 7 points, the odd-vs-even and neutral-midpoint debates, when each fits, and how AI follow-ups make any scale richer.