{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-12T14:04:27.791Z"},"content":[{"type":"documentation","id":"6783881a-91ea-41d7-a48b-b3e4b9378ab8","slug":"futility-analysis-when-to-stop-a-study","title":"Futility Analysis: How to Decide a Running Study Will Never Answer Your Question","url":"https://www.koji.so/docs/futility-analysis-when-to-stop-a-study","summary":"There are three reasons to stop a study early: efficacy, futility, and harm. Stopping for benefit is the dangerous one - Montori et al. (JAMA 2005) found 143 truncated trials recruited 63% of planned sample, stopped at a median of 66 events, reported a median risk ratio of 0.53, and 94% failed to report at least one key methodological detail. Bassler et al. (JAMA 2010) found truncated trials overstate effects (pooled ratio of relative risks 0.71) independent of whether a stopping rule existed, and in 62% of questions the non-truncated evidence showed no significant benefit. Futility is safer because you decline to make a claim; practical triggers are recruitment, variance, question and decision futility. Stop-for-harm is exemplified by CAST (NEJM 1989), where the drugs suppressed the surrogate endpoint and raised total mortality from 3.0% to 7.7%. The FDA 2024 DMC draft guidance requires monitoring bodies independent of the sponsor and trial conduct - the transferable principle is that the person who wants the result should not be the person who decides to stop.","content":"**Answer first: there are exactly three reasons to stop a study before its planned end - it worked, it will never work, or it is causing harm. Product teams routinely act on the first, almost never plan for the second, and have no process at all for the third.** That ordering is backwards. The evidence from clinical trials is that stopping early for benefit produces systematically overstated effects: in a systematic review of 91 truncated trials matched against 424 trials that ran to completion, the pooled ratio of relative risks was 0.71, and in 39 of the 63 clinical questions studied (62%) the full non-truncated evidence base failed to show a significant benefit at all. Stopping early because it will never work, by contrast, costs nothing but the budget you save, and it is the decision nobody schedules.\n\nThis guide covers futility rules, the stop-for-harm case that product research has no vocabulary for, and the governance question underneath both: who gets to call it.\n\n## The three stopping reasons, and how differently they behave\n\nClinical trials are monitored on all three continuously. The asymmetry between them is the single most useful thing to import.\n\n| Reason to stop | What it claims | Evidence quality when you act on it | How often product teams plan for it |\n| --- | --- | --- | --- |\n| Efficacy | The effect is real and large enough to act now | Systematically overstated, worst with few events | Constantly, informally |\n| Futility | The study will not reach a usable answer | Robust - you are declining to claim anything | Almost never |\n| Harm | The thing being tested is hurting people | Usually decisive | No process at all |\n\n### Stopping for success is the dangerous one\n\nTwo systematic reviews establish this and they are worth knowing by number.\n\nMontori and colleagues (JAMA, 2005, 294(17):2203-2209) identified **143 randomized trials stopped early for benefit**, 92 of them published in five high-impact medical journals. The proportion of trials in those journals stopped early for benefit rose from 0.5% in 1990-1994 to 1.2% in 2000-2004. On average these trials recruited **63% of their planned sample** and stopped after a median of 13 months of follow-up, one interim analysis, and a median of just 66 events. The median risk ratio was 0.53 - an apparent halving of risk. And **135 of the 143 (94%) failed to report at least one of**: the planned sample size, the interim analysis after which the trial stopped, whether a stopping rule informed the decision, or an adjusted analysis accounting for truncation. That last item was missing in 129 of 143. Trials with fewer events reported larger effects, with an odds ratio of 28 (95% CI 11 to 73).\n\nBassler and colleagues (JAMA, 2010, 303(12):1180-1187) then compared 91 truncated trials against 424 matched trials that were not stopped early. The pooled ratio of relative risks was **0.71 (95% CI 0.65 to 0.77)** - truncated trials reported effects about 29% larger. Critically, **this difference was independent of whether a statistical stopping rule was present** and independent of methodological quality. The overstatement was worst in trials with fewer than 500 events. And in **39 of 63 questions (62%), the pooled effect from the non-truncated trials failed to demonstrate significant benefit** at all.\n\nThe translation to product research is direct. A study stopped at 40% of target because the early numbers looked strong is the research equivalent of a truncated trial, with far fewer observations than 500 events and no adjustment. Whatever effect size it reports should be treated as an upper bound, not an estimate. If you must stop early for success, do it against a pre-declared boundary - see [interim analysis and sequential testing](/docs/interim-analysis-sequential-testing-research) for the thresholds - and report the effect as provisional.\n\n## Futility: the decision that saves the most money\n\nFutility asks a different question. Not \"is the effect real\" but \"given what we have so far, is there any plausible way the remaining sample changes the conclusion?\"\n\nThe standard tool is **conditional power**: the probability that the study reaches its threshold at the planned end, given the data already collected and an assumption about the true effect. If conditional power is very low - the conventional trigger is somewhere in the 10% to 20% range - continuing is spending budget to confirm something you already know.\n\nFutility has a property that makes it unusually safe to act on: **you are declining to make a claim, not making one.** The overstatement problem that afflicts efficacy stopping does not apply, because there is no effect estimate being published. The worst case is that you abandoned a study that would have squeaked over the line, which is a cost, not an error.\n\n### Futility rules that work without conditional power\n\nMost customer research cannot compute conditional power, and does not need to. Futility in practice is usually structural rather than statistical, and these four rules cover the great majority of wasted studies:\n\n1. **Recruitment futility.** The study needs 120 responses from a segment that has produced 4 in two weeks. No analysis will fix a sample that does not exist. Set the rule as a rate: \"if we are below 40% of target at the halfway date, we stop and redesign recruitment.\"\n2. **Variance futility.** The outcome is so noisy that the confidence interval at full sample would still contain both \"large improvement\" and \"no change.\" This is a power problem discovered late; see [statistical power and minimum detectable effect](/docs/statistical-power-minimum-detectable-effect) for how to catch it before fielding.\n3. **Question futility.** Respondents are not answering the question you asked. If the open_ended responses to your key item are consistently about something else, more of them will not help. The fix is a new instrument, not a bigger sample.\n4. **Decision futility.** The most underrated one. Ask, at the halfway look: if the result comes back at the most favourable plausible value, does anyone change what they were going to do? If not, the study is futile regardless of its statistics, and it was futile before it launched. Our guide to [research peer review](/docs/research-peer-review-qa-gate) covers catching this at the pre-launch gate, which is where it belongs.\n\n## Stopping for harm: the case product research has no words for\n\nThe definitive example is the Cardiac Arrhythmia Suppression Trial (CAST), reported in the New England Journal of Medicine in 1989 (321(6):406-412). The premise was sound: ventricular premature depolarizations after a heart attack predict sudden death, so suppressing them should save lives. Of 2,309 patients recruited to the titration phase, 1,727 (75%) had their arrhythmia successfully suppressed by one of the study drugs and were randomized to active drug or placebo.\n\nThe drugs worked. The surrogate endpoint moved exactly as intended. Over an average of 10 months, patients on encainide or flecainide had **33 arrhythmic deaths or cardiac arrests out of 730 (4.5%), against 9 out of 725 on placebo (1.2%) - a relative risk of 3.6 (95% CI 1.7 to 8.5)**. Total mortality was **56 of 730 (7.7%) versus 22 of 725 (3.0%), a relative risk of 2.5 (95% CI 1.6 to 4.5)**. That arm of the trial was stopped.\n\nCAST is the canonical demonstration that **a surrogate metric can move in the right direction while the outcome you actually care about moves in the wrong one.** Every product team that optimises engagement while retention quietly falls, or reduces support tickets by making the help page harder to find, is running a small CAST. The reason clinical trials catch it and product research does not is not superior ethics; it is that somebody is explicitly tasked with watching the harm endpoint, and that person is not the one who wants the treatment to work.\n\nIn research specifically, harm has a narrower and very real meaning: a study can hurt its participants. Questions that surface trauma, screeners that expose sensitive status, an incentive structure that pressures people into completing, or an interview that runs three times its advertised length. A harm-stopping rule for a research study is a sentence in the brief naming the signal that halts fielding immediately - a distress report, a complaint pattern, an unexpected drop-off cliff on a sensitive item.\n\n## Who is allowed to call it\n\nThis is the part that transfers most cleanly and that almost nobody has implemented.\n\nClinical trials use a Data Monitoring Committee (also called a DSMB or DMC). It is typically a committee of three to nine external experts - clinicians, one or two statisticians, and often an ethicist - who meet once or twice a year, review trial conduct, safety, futility and efficacy, and recommend to the sponsor whether to continue, modify or stop. The sponsor makes the final decision but in practice almost always accepts the recommendation.\n\nThe FDA draft guidance *Use of Data Monitoring Committees in Clinical Trials* (February 2024) states the structural principle plainly: a DMC \"is established by the sponsor but should be independent of the sponsor and the trial conduct.\" And on the specific decision this guide is about: \"Changes to the trial design that involve an analysis of results by study group are best performed by a body independent of the sponsor, the investigators, and the subjects.\"\n\nNote what that sentence rules out. The person who proposed the feature, the person running the study, and the person who will present the results should not be the person deciding, on the basis of interim results, whether to stop. In most product teams that is one person, and it is usually the same person who is watching the dashboard every morning.\n\n### A lightweight version that fits a product org\n\nYou do not need a nine-person committee for a 60-participant interview study. You need the firewall, not the ceremony.\n\n| Element | Clinical version | Product research version |\n| --- | --- | --- |\n| Who monitors | External DMC, 3-9 members | One named reviewer outside the requesting team |\n| What they see | Unblinded comparative interim data | Comparative interim results, which nobody else sees |\n| What everyone else sees | Aggregate, blinded reports | Operational metrics only - pace, completion, segment fill |\n| Decision authority | Recommends to sponsor | Recommends to the study owner, in writing |\n| Governing document | DMC charter | Three lines in the study brief |\n| Conflict rule | No ongoing financial relationship with sponsor | Reviewer does not own the roadmap item being tested |\n\nThe three lines in the brief are: who may see comparative interim results, what the futility and harm triggers are, and who signs off on stopping. That is the whole intervention. It costs nothing and it removes the single most common way a research programme talks itself into a conclusion.\n\n## How Koji supports stopping decisions\n\n**Operational monitoring is separable from outcome analysis.** Koji reports show recruitment pace, completion status and per-question drop-off alongside findings, so the reviewer watching for recruitment futility does not have to read the substantive results to do their job. That separation is what makes a firewall practical rather than theoretical.\n\n**Parallel fielding shortens the window in which futility is expensive.** Because AI-moderated voice and text interviews run concurrently rather than one at a time, a study that would have taken three weeks of sequential scheduling completes in days. Futility discovered on day two costs two days. The same futility discovered in week three of manual scheduling costs three weeks and most of the budget.\n\n**Structured questions make futility triggers checkable.** With all six question types available - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - the primary outcome can be a specific item with a defined distribution, so a futility rule like \"if fewer than 15% select this option at halfway, we stop\" is something a monitor can evaluate without interpretation. See the [structured questions guide](/docs/structured-questions-guide). A study whose only outcome is free text has no checkable futility trigger, which is one reason free-text-only studies run long.\n\n**Credits follow completed quality conversations.** Koji applies a quality gate so that only conversations meeting a quality threshold consume credits, which means a study stopped for futility does not bill for the low-quality tail it was collecting on its way to nowhere.\n\nCompared with traditional survey platforms such as SurveyMonkey, Typeform or Qualtrics - where a fielded survey simply runs until its close date and the only monitoring surface is a response counter - platforms like Koji give a monitor enough structure to act on a stopping rule while the study is still cheap to stop.\n\n## Frequently asked questions\n\n### What is the difference between futility and just not having enough data yet?\n\nFutility is a forward-looking judgment: given what you have, the remaining planned sample cannot plausibly change the conclusion. Not having enough data yet means the remaining sample can change it and you should keep going. The test is arithmetic in quantitative work (conditional power) and structural in qualitative work (can the recruitment, the instrument or the decision context deliver an answer at all).\n\n### Is stopping a study early for good results ever acceptable?\n\nYes, when the boundary was set in advance, the study crosses it, and the readout reports the effect as an upper bound rather than a point estimate. The evidence is unambiguous that truncated studies overstate effects - by roughly 29% on average in the Bassler analysis - and that the overstatement persists even when a formal stopping rule was in place. Treat an early win as a reason to plan a confirmatory study, not as the confirmation.\n\n### Who should sit on a monitoring role in a small company?\n\nSomeone who does not own the outcome. A researcher from an adjacent team, a data analyst, or the person who ran the pre-launch peer review. The only hard requirement is that stopping the study early, in either direction, must not affect anything they are accountable for.\n\n### How do you set a futility rule for a qualitative study?\n\nUse recruitment and instrument criteria rather than statistical ones. For example: \"if we have not completed 12 interviews in the primary segment by day 7, we stop and revisit screening\" or \"if three consecutive participants cannot answer the core question as written, we stop and rewrite it.\" Both are checkable, and both catch the two ways qualitative studies actually fail.\n\n### What is a harm-stopping rule in customer research?\n\nA named signal that halts fielding immediately, written into the brief before launch. Common ones: any participant reporting distress, a complaint pattern about a specific question, an unexpected drop-off cliff on a sensitive item, or discovery that a screener is exposing information participants did not consent to share. The point of writing it down is that in the moment, the person who sees the signal is usually the person least empowered to stop the study.\n\n### Does this apply to always-on or continuous research programmes?\n\nYes, and futility matters more there, because a continuous programme has no natural end date to force the question. Schedule a standing futility review - quarterly is typical - that asks whether the programme is still changing decisions. Decision futility is the failure mode of continuous discovery, and it is invisible without a scheduled check.\n\n## Related Resources\n\n- [Interim Analysis and Sequential Testing](/docs/interim-analysis-sequential-testing-research) - the boundaries that make stopping for success defensible\n- [Intention to Treat vs Per Protocol](/docs/intention-to-treat-per-protocol-research) - which responses count once you have stopped\n- [Changing a Study While It Is Running](/docs/changing-a-study-mid-field) - the alternative to stopping, and its own rules\n- [Research Peer Review: The Pre-Launch QA Gate](/docs/research-peer-review-qa-gate) - catching decision futility before you spend anything\n- [Statistical Power and Minimum Detectable Effect](/docs/statistical-power-minimum-detectable-effect) - sizing a study so futility is not the default outcome\n- [Publication Bias in Product Research](/docs/publication-bias-product-research) - what happens to the studies that get stopped and never written up\n- [Structured Questions Guide](/docs/structured-questions-guide) - defining outcomes a monitor can actually check","category":"Research Methods","lastModified":"2026-08-12T03:26:05.425871+00:00","metaTitle":"Futility Analysis: When to Stop a Research Study Early (2026)","metaDescription":"Trials stopped early for benefit overstate effects by ~29%, and in 62% of questions the full evidence showed no benefit at all. Learn futility rules, stop-for-harm triggers, and who should make the call.","keywords":["futility analysis","stopping rules","when to stop a study early","conditional power","data monitoring committee","stop for harm","truncated trials","research operations"],"aiSummary":"There are three reasons to stop a study early: efficacy, futility, and harm. Stopping for benefit is the dangerous one - Montori et al. (JAMA 2005) found 143 truncated trials recruited 63% of planned sample, stopped at a median of 66 events, reported a median risk ratio of 0.53, and 94% failed to report at least one key methodological detail. Bassler et al. (JAMA 2010) found truncated trials overstate effects (pooled ratio of relative risks 0.71) independent of whether a stopping rule existed, and in 62% of questions the non-truncated evidence showed no significant benefit. Futility is safer because you decline to make a claim; practical triggers are recruitment, variance, question and decision futility. Stop-for-harm is exemplified by CAST (NEJM 1989), where the drugs suppressed the surrogate endpoint and raised total mortality from 3.0% to 7.7%. The FDA 2024 DMC draft guidance requires monitoring bodies independent of the sponsor and trial conduct - the transferable principle is that the person who wants the result should not be the person who decides to stop.","aiPrerequisites":["Familiarity with running a study end to end","Basic understanding of confidence intervals and effect sizes"],"aiLearningOutcomes":["Distinguish stopping for efficacy, futility and harm, and treat them differently","Write recruitment, variance, question and decision futility triggers into a brief","Recognise a surrogate-endpoint trap of the CAST type in product metrics","Set up a lightweight monitoring firewall that fits a product organisation"],"aiDifficulty":"intermediate","aiEstimatedTime":"12 min"}],"pagination":{"total":1,"returned":1,"offset":0}}