{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-12T15:07:24.771Z"},"content":[{"type":"documentation","id":"6444b863-6600-4564-9ca9-4e54696468d7","slug":"changing-a-study-mid-field","title":"Changing a Study While It Is Running: Pre-Planned Adaptations vs Protocol Amendments","url":"https://www.koji.so/docs/changing-a-study-mid-field","summary":"A change to a running study is either a prospectively planned adaptation or a reactive amendment. The FDA 2019 adaptive designs guidance defines an adaptive design as allowing prospectively planned modifications based on accumulating data, where prospective means planned and specified before any comparative analyses are conducted. Three kinds of change: blind or non-comparative changes (typo fixes, screener corrections, pooled-variance sample size re-estimation) cost nothing; pre-planned adaptations (adding a segment, dropping a concept, extending fielding) are legitimate only if the rule was written first; reactive amendments made after seeing comparative results split the study into two and must not be pooled. The root-cause control is limiting access to comparative interim results to people independent of trial conduct, because once you have seen results you cannot prove a change was not motivated by them. Keep a six-column amendment log. For AI-moderated research, adaptation in the probe (follow-ups varying under a fixed instruction set) is not adaptation in the design (the rules themselves changing).","content":"**Answer first: a change made to a running study is either an adaptation you planned before you saw any comparative results, or an amendment you made because of what you saw. The first is a design feature with known statistical properties. The second is a new study wearing the old study's sample.** The FDA draws exactly this line. Its 2019 guidance *Adaptive Designs for Clinical Trials of Drugs and Biologics* defines an adaptive design as \"a clinical trial design that allows for prospectively planned modifications to one or more aspects of the design based on accumulating data from subjects in the trial,\" and specifies that prospective \"means that the adaptation is planned and details specified before any comparative analyses of accumulating trial data are conducted.\"\n\nThis is the capstone of the running-study cluster, because every discipline in the other three guides assumes the instrument holds still. [Interim analysis](/docs/interim-analysis-sequential-testing-research) prices your looks against a fixed primary outcome. [Futility rules](/docs/futility-analysis-when-to-stop-a-study) assume the thing being judged futile is the thing you started. [Analysis populations](/docs/intention-to-treat-per-protocol-research) assume everyone answered the same questions. Change the study mid-field and all three quietly stop applying - and unlike stopping early, nobody flags it, because changing a question feels like an improvement rather than a decision.\n\n## The three kinds of mid-study change\n\n| Kind | Trigger | Statistical cost | What to do |\n| --- | --- | --- | --- |\n| Pre-planned adaptation | A rule written in the brief fires | Priced into the design | Execute it and log that it fired |\n| Blind or non-comparative change | Something you can see without looking at results by group | Usually none | Make it, log it, note the date |\n| Reactive amendment | You saw comparative results and responded | Uncontrolled and unquantifiable | Treat the study as split into two |\n\nThe distinction that does all the work is not *what* changed. It is *what you knew when you changed it*.\n\n### Changes you can make freely, because they are blind to the answer\n\nThese depend on information that does not involve comparing groups on the outcome, so they cannot be motivated by the result:\n\n- Fixing a typo or a broken option in a single_choice list.\n- Adding a screener criterion because you discovered the recruitment source is contaminated with the wrong population.\n- Extending fielding because response pace is below target.\n- Adjusting an incentive because completion rate is low.\n- Re-estimating sample size from pooled variance without splitting by group - the FDA guidance treats this as an adaptation based on non-comparative data, which is why it is the least controversial adaptation there is.\n\nThe test to apply: **could the change have gone the other way if the results had been reversed?** If a stranger who has seen no outcome data would make the same change, it is blind. Log it and continue.\n\n### Changes that need a pre-planned rule\n\nThese are legitimate adaptations, but only if the rule was written before any comparative look. In clinical trials the standard examples are sample size re-estimation, dropping a treatment arm, adaptive population selection and adaptive randomization. In research the equivalents are:\n\n- **Adding a segment** when a pre-specified subgroup fills below target.\n- **Dropping a concept** from a multi-concept test when it fails a pre-declared screening threshold.\n- **Extending or truncating fielding** against a pre-declared futility or efficacy rule.\n- **Escalating probing depth** on a question when a pre-declared proportion of responses come back uninformative.\n\nThe FDA guidance explains precisely why the rule has to be written in advance rather than merely applied honestly: complete prespecification \"helps increase confidence that adaptation decisions were not based on accumulating knowledge in an unplanned way,\" and specifying the exact rule \"reduces concern that the adaptation could have been influenced by knowledge of comparative results.\" Good faith is not the mechanism. The written rule is.\n\n### Changes that split the study in two\n\nRewording the primary question. Changing the scale from 1-5 to 1-7. Swapping a yes_no item for a multiple_choice one. Adding a leading follow-up because early responses hinted at something. Restricting analysis to a subgroup that looked promising at the halfway mark.\n\nThe FDA guidance describes this scenario exactly: when data are examined in a comparative interim analysis, analyses that were not prospectively planned \"may unexpectedly appear to indicate that some specific design change (e.g., restricting analyses to some population subset, dropping a treatment arm, adjusting sample size, modifying the primary endpoint, or changing analysis methods) is ethically important or might increase the potential for a statistically significant final trial result.\" The consequence: \"Such revisions based on non-prospectively planned analyses can create difficulty in controlling the Type I error probability and in interpreting the trial results. Sponsors are strongly discouraged from implementing such changes without first meeting with FDA.\"\n\nYou do not have a regulator to meet with. What you have instead is a reporting obligation. **If the instrument changed on the basis of comparative results, the honest analysis treats pre-change and post-change respondents as two studies.** You can report both. You cannot pool them and describe the pooled number as one measurement, because the two halves were not measuring the same thing on the same population under the same conditions.\n\n## The integrity control that prevents most of this\n\nReactive amendments are almost always downstream of one root cause: the people who could change the study could also see the comparative results.\n\nThe FDA states the control directly: \"it is strongly recommended that access to comparative interim results be limited to individuals with relevant expertise who are independent of the personnel involved in conducting or managing the trial and have a need to know.\" The reason given is worth quoting because it is not about honesty: knowledge of comparative interim results by trial management personnel \"may make it difficult for regulators to determine whether a protocol amendment seemingly well-motivated by information external to\" the trial actually was. In other words, once you have seen the results, **you can no longer prove your reasonable change was not motivated by them - and neither can anyone else, including you.**\n\nThe practical version for a product team is three rules:\n\n1. The person who can edit a live study does not watch the comparative results.\n2. Anyone who does watch comparative results has no edit rights while fielding is open.\n3. If the roles must be combined - and in a small team they often must - every edit made after the first look is logged with a timestamp and a stated reason, and the readout reports it.\n\nThat third rule is the realistic one for most companies, and it is worth more than it looks. An amendment log turns an invisible problem into a visible one, and visible problems get argued about on the merits.\n\n## The amendment log\n\nSix columns. Keep it with the brief, not in a chat thread.\n\n| Column | Why it is there |\n| --- | --- |\n| Date and time | Establishes whether the change preceded or followed a look |\n| What changed | The specific question, option, screener or rule |\n| Trigger | The pre-planned rule that fired, or the observation that prompted it |\n| Comparative results seen first? | Yes or no. This one column determines everything else |\n| Respondents before / after | The two sub-samples the analysis may have to separate |\n| Analysis consequence | Pooled, split, or excluded |\n\nA study with an empty amendment log and a study with no amendment log look identical in the readout and are completely different objects. Publish the empty one.\n\n## The AI-moderated case: adaptation in the probe is not adaptation in the design\n\nThis is the point where an honest guide has to turn the argument on its own tooling. An AI moderator that generates follow-up questions is, by any plain reading, changing the instrument during the study. Every participant gets a different conversation. So does the whole framework collapse?\n\nNo, but only because of a distinction that has to be maintained deliberately.\n\n**Adaptation in the probe** means the follow-up varies with what the participant just said, under a fixed instruction set that was written before fielding. Participant 3 and participant 30 face the same questions, the same probing rules and the same depth limits; what differs is their own answers. This is what a good human interviewer does, and it is the reason interviews outperform static surveys.\n\n**Adaptation in the design** means the question set, the primary outcome, the probing instructions or the eligibility rules change between participants because of what earlier participants said. That is a protocol amendment, and it is the thing the FDA definition rules out unless it was pre-planned.\n\nThe line between them is whether the *rule* changed or only the *realisation of the rule*. And the practical consequence is reassuring: an AI moderator applying a fixed instruction set is more protocol-compliant than a human moderator, not less, because the human drifts. A human running twenty interviews across three weeks probes harder on whatever seemed promising in week one - an unlogged, unnoticed protocol amendment that happens in every manual study and appears in none of their limitations sections. See [model version drift](/docs/ai-model-version-drift-research) for the one genuine AI-specific version of this problem, which is the underlying model changing beneath a study that has not changed at all.\n\n## How Koji keeps the design fixed while the conversation adapts\n\n**Structured questions are the fixed frame.** The six question types - open_ended, scale, single_choice, multiple_choice, ranking and yes_no - define the instrument. A scale question asked at position 4 with the same anchors is the same measurement for participant 3 and participant 300, whatever the AI probed around it. That is what makes cross-participant aggregation legitimate at all. See the [structured questions guide](/docs/structured-questions-guide).\n\n**Probing instructions are set at design time, per question.** Follow-up depth and probing guidance are configured before fielding rather than adjusted live, so the adaptive behaviour is a property of the design, not a running intervention.\n\n**Branching is pre-declared logic, not a mid-study edit.** Conditional paths through an interview are a pre-planned adaptation in the exact FDA sense: the rule, the triggering data and the consequence are all specified before anyone is fielded. See [adaptive AI interview branching](/docs/adaptive-ai-interview-branching).\n\n**Every conversation is retained with its status and transcript**, so if an amendment does happen you can split the analysis cleanly at the change point rather than discovering that the pre-change responses were overwritten.\n\nTraditional survey tools such as SurveyMonkey, Typeform and Qualtrics let you edit a live survey with no record, no warning and no way to reconstruct which respondents saw which version. Platforms like Koji do not make the underlying judgment for you, but a fixed structured frame plus a full conversation record is the difference between an amendment you can analyse around and one you can only apologise for.\n\n## Frequently asked questions\n\n### Can I fix a typo in a live study?\n\nYes. A typo fix is blind to the outcome - you would make it whichever way the results were going - so it costs nothing statistically. Log it with a timestamp anyway, because the log is what lets you prove it was that kind of change rather than the other kind.\n\n### What if a question is clearly confusing people mid-field?\n\nIf you noticed from operational signals - drop-off at that question, response times, participants saying they do not understand - that is a non-comparative observation and you can act on it. Fix the question, log it, and treat responses before and after as two versions of that item. Do not pool them into a single distribution. If you noticed because the answers were not supporting your hypothesis, that is a reactive amendment and a much bigger problem.\n\n### Does adding a new segment mid-study invalidate the results?\n\nNot for the original segment, whose data is unaffected. The new segment is a later, smaller, separately recruited sample and should be reported as such, with its own fielding dates. The failure mode is pooling them and reporting one overall number, because the two groups were fielded under different conditions at different times.\n\n### How is this different from an adaptive AI interview?\n\nAn AI moderator varies the follow-up questions under a fixed instruction set written before fielding, which is adaptation in the probe. A protocol amendment changes the question set, the primary outcome or the probing rules themselves. The test is whether the rule changed or only its realisation for a particular participant.\n\n### Do I need to disclose mid-study changes in the readout?\n\nYes, in one line: what changed, when, and whether comparative results had been seen first. Reviewers now assume undisclosed changes happened, so disclosure is a credibility gain rather than a cost - and a study that reports an empty amendment log is making a stronger claim than one that says nothing.\n\n### What about changes forced by something outside the study?\n\nThe FDA guidance handles this case explicitly: unpredictable external events - new safety information, a competitor launch, a pricing change - may justify revising a running study rather than terminating it and starting over. The same rules apply. Log the external trigger, note the date, and split the analysis if the instrument changed. An external motivation is easier to defend precisely because it is documentable independently of your results.\n\n## Related Resources\n\n- [Interim Analysis and Sequential Testing](/docs/interim-analysis-sequential-testing-research) - what a look costs, and why seeing results changes what you may do next\n- [Futility Analysis](/docs/futility-analysis-when-to-stop-a-study) - stopping as the alternative to amending\n- [Intention to Treat vs Per Protocol](/docs/intention-to-treat-per-protocol-research) - what an amendment does to your analysis population\n- [Model Version Drift in Research](/docs/ai-model-version-drift-research) - when the instrument changes without anyone editing it\n- [Adaptive AI Interview Branching](/docs/adaptive-ai-interview-branching) - pre-declared conditional logic done properly\n- [Research Peer Review: The Pre-Launch QA Gate](/docs/research-peer-review-qa-gate) - the cheapest place to catch a change you would otherwise make mid-field\n- [Structured Questions Guide](/docs/structured-questions-guide) - the fixed frame that lets conversations vary safely","category":"Research Operations","lastModified":"2026-08-12T03:26:05.425871+00:00","metaTitle":"Changing a Study Mid-Field: Adaptations vs Protocol Amendments (2026)","metaDescription":"A mid-study change is either a pre-planned adaptation with known properties or a reactive amendment that splits your study in two. Learn the FDA distinction, the integrity control, and the amendment log.","keywords":["protocol amendment","adaptive design","mid-study changes","changing a survey while live","sample size re-estimation","research operations","study integrity","pre-planned adaptation"],"aiSummary":"A change to a running study is either a prospectively planned adaptation or a reactive amendment. The FDA 2019 adaptive designs guidance defines an adaptive design as allowing prospectively planned modifications based on accumulating data, where prospective means planned and specified before any comparative analyses are conducted. Three kinds of change: blind or non-comparative changes (typo fixes, screener corrections, pooled-variance sample size re-estimation) cost nothing; pre-planned adaptations (adding a segment, dropping a concept, extending fielding) are legitimate only if the rule was written first; reactive amendments made after seeing comparative results split the study into two and must not be pooled. The root-cause control is limiting access to comparative interim results to people independent of trial conduct, because once you have seen results you cannot prove a change was not motivated by them. Keep a six-column amendment log. For AI-moderated research, adaptation in the probe (follow-ups varying under a fixed instruction set) is not adaptation in the design (the rules themselves changing).","aiPrerequisites":["Experience fielding a study end to end","Familiarity with interim analysis and stopping rules"],"aiLearningOutcomes":["Classify any mid-study change as blind, pre-planned or reactive","Apply the FDA prospective-planning test to your own study edits","Keep an amendment log that makes changes auditable","Distinguish adaptive AI probing from a protocol amendment"],"aiDifficulty":"advanced","aiEstimatedTime":"11 min"}],"pagination":{"total":1,"returned":1,"offset":0}}