{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-14T17:20:25.364Z"},"content":[{"type":"documentation","id":"8bcb4b26-903e-4700-a8dc-6e1463a64f5e","slug":"stepped-wedge-rollout-research","title":"The Stepped Wedge: How to Randomise a Rollout You Cannot Randomise (2026)","url":"https://www.koji.so/docs/stepped-wedge-rollout-research","summary":"A stepped wedge cluster randomised trial starts every cluster in the control condition and crosses them over to the intervention in a randomly assigned order until all are exposed. It solves the case where a change must reach everyone, so who receives it cannot be randomised but when they receive it can. Originated in the Gambia Hepatitis Intervention Study from 1986, where 17 vaccination teams were randomly assigned starting dates. It suits B2B software because features roll out per account and rollouts are already phased. The main risk is confounding with time, which requires modelling period explicitly; reviews show reporting of time handling is often inadequate.","content":"**The standard advice for a change that ships to everyone is to give up on randomisation and reconstruct a counterfactual afterwards. That advice is often wrong, because it answers the wrong question. You cannot randomise WHO receives the change. You can almost always randomise WHEN they receive it, and that is enough to get a randomised comparison.**\n\nThis is the stepped wedge cluster randomised trial. Every unit starts in the control condition, units cross over to the intervention in a randomly assigned order, and by the end everyone has it. Nobody is denied anything. The rollout you were going to do anyway becomes the experiment.\n\nIf you have already read [quasi-experimental design](/docs/quasi-experimental-design-guide), this is the design that rescues a case that article correctly concedes. Quasi-experimental methods construct a counterfactual from pre-existing trends or untreated comparison groups because assignment was not random. The stepped wedge does not need to construct one, because the assignment was random after all. **The difference between the two is not statistical sophistication. It is whether somebody thought about the rollout order before the rollout started.**\n\n## The design\n\nHemming and colleagues describe it in the *BMJ* (2015;350:h391) as involving \"random and sequential crossover of clusters from control to intervention until all clusters are exposed\". Every cluster provides both before and after observations, and every cluster switches, but not at the same time.\n\nA four-step rollout across twelve accounts looks like this, where C is control and I is intervention:\n\n| Group | Period 1 | Period 2 | Period 3 | Period 4 | Period 5 |\n| --- | --- | --- | --- | --- | --- |\n| Accounts 1-3 | C | I | I | I | I |\n| Accounts 4-6 | C | C | I | I | I |\n| Accounts 7-9 | C | C | C | I | I |\n| Accounts 10-12 | C | C | C | C | I |\n\nThe group assignment to rows is the randomised part. Reading down any column gives you a contemporaneous comparison between accounts that have the feature and accounts that do not. Reading across any row gives you a within-account before-and-after. **The design supplies both kinds of comparison from a rollout that most teams would have sequenced by account size or by whoever asked loudest.**\n\n## Where it came from\n\nThe first stepped wedge trial was the Gambia Hepatitis Intervention Study, launched in July 1986. The goal was to evaluate adding hepatitis B vaccine to the existing childhood immunisation schedule. Instantaneous nationwide vaccination was impossible for logistical and financial reasons, so the 17 teams delivering the existing vaccines were **randomly assigned starting dates** for adding the new one, phased over about four years.\n\nThe reasoning translates to product work almost word for word. The rollout had to be phased because of capacity. Since it had to be phased, the order had to be decided somehow. Deciding it at random cost nothing and produced an unvaccinated comparison group at every point in time, purely as a by-product.\n\n## Why this fits B2B software better than it fits medicine\n\nStepped wedge designs are used in health services research because interventions there are often delivered to whole hospitals or clinics rather than to individuals. That is exactly the shape of B2B software.\n\n- Features are frequently enabled per workspace, per account or per region, not per user.\n- Rollouts are already phased, because support, training and infrastructure capacity are finite.\n- Everyone is contractually or politically entitled to the feature eventually, so withholding it from a control group indefinitely is not an option.\n- The unit that experiences the change is the team, and outcomes such as retention and expansion are measured at account level anyway.\n\nEvery one of those facts is usually cited as a reason experimentation is impossible. **They are the exact conditions the stepped wedge was invented for.**\n\n| Objection to experimenting | What the stepped wedge does with it |\n| --- | --- |\n| \"We cannot withhold it from customers\" | Nothing is withheld. Everyone receives it, in a random order |\n| \"It ships to the whole account, not to users\" | The account is the cluster. That is the intended unit |\n| \"We only have 15 enterprise accounts\" | Each account contributes both control and intervention periods, so the design is unusually efficient at small cluster counts |\n| \"The rollout is phased for support capacity\" | Then an order already exists. Randomising it is free |\n| \"Leadership will not wait for a trial\" | The trial finishes when the rollout finishes. It adds no delay |\n\nThe small-sample point deserves emphasis because it is the most common reason B2B teams believe experimentation is closed to them. In a parallel design, 15 accounts split into two groups is a hopeless comparison. In a stepped wedge, each account serves as its own control before it crosses over, and the design accumulates comparisons at every step.\n\n## The cost: time is now a confounder\n\nThe stepped wedge is not free, and the price is specific. Because the intervention is perfectly correlated with time, **anything else that changes over the rollout window is confounded with the feature.** A seasonal uplift, a pricing change, a competitor exit or a support reorganisation will all masquerade as the effect if the analysis ignores time.\n\nThis is why the design requires modelling the underlying temporal trend rather than simply comparing before against after. Hemming and Lilford note that results can \"transpire to be the result of a positive underlying temporal trend\", and that a large number of published studies do not report how they allowed for time in the design or the analysis.\n\nThe reporting record is genuinely poor. In a review of 123 studies using the design, of which 39 were completed trial reports, Grayling, Wason and Mander found that reporting quality on individual criteria fell as low as 15.4 percent, with a median of 66.7 percent (*Trials* 2017;18:33). Only 25 of the 39 completed trials, or 64.1 percent, gave a rationale for choosing the stepped wedge at all. The median trial used 20.5 clusters and 9 steps. The CONSORT reporting extension for these trials (Hemming, Taljaard and colleagues, *BMJ* 2018;363:k1614) exists partly because of this, and it requires investigators to justify the design choice explicitly.\n\nThe lesson for a product team is to decide three things in writing before the first account crosses over:\n\n1. **How many steps, and how long is each?** A step must be long enough for the outcome to be observable within it.\n2. **How will time be modelled?** At minimum, period should enter the analysis as a term. Without it the estimate is not interpretable.\n3. **What is the lag before the effect is expected?** If a feature takes six weeks to change behaviour, a four-week step attributes the change to the wrong period.\n\n## What to measure at each step, and why teams abandon this\n\nThe design's weak point in practice is not the statistics. It is that a stepped wedge needs measurement from every cluster in every period, including the periods where nothing has happened yet. Those control-period observations are what the whole design rests on.\n\nBehavioural telemetry covers part of it automatically. The part teams give up on is the human evidence, because interviewing a sample of accounts at five separate points across a rollout is, with traditional methods, five recruitment cycles, five scheduling rounds and five analysis passes. Most teams run interviews once, at the end, with the accounts that are happiest to talk, which reintroduces exactly the [conditioning problem](/docs/collider-bias-product-research) the design was supposed to avoid.\n\nThis is where the economics have genuinely changed. Koji makes the repeated qualitative arm affordable:\n\n- **AI-moderated interviews can run at every step across every cluster**, including the accounts still in the control condition, without consuming a researcher week per wave.\n- **Identical moderation across periods and clusters.** This is not a convenience point. If wave 5 is probed more thoroughly than wave 1 because the moderator has learned what is interesting, the apparent effect includes the moderator's learning curve. See [surveillance bias](/docs/surveillance-bias-detection-research) for how much damage unequal probing does.\n- **Voice interviews** reach operational users inside customer accounts who will never join a scheduled call.\n- **Automatic thematic analysis** makes it feasible to compare open-ended themes across ten cells rather than reading them once at the end.\n- **Real-time reporting** means a step can inform the next step, which is the entire point of a phased rollout.\n\nThe six structured question types carry unusual weight in this design. Koji supports `open_ended`, `scale`, `single_choice`, `multiple_choice`, `ranking` and `yes_no`. Because a stepped wedge compares across time, **you need items whose meaning does not drift between period 1 and period 5**, and closed types are the ones that hold still. A fixed `scale` item on workflow confidence, a `yes_no` on whether a workaround is still in use, a `single_choice` on the primary tool for the task, and a `ranking` of priorities, all repeated identically at every step, produce a clean time series alongside the telemetry. The `open_ended` responses then explain the movements rather than having to detect them. The [structured questions guide](/docs/structured-questions-guide) covers building an instrument that survives repetition.\n\n## When not to use it\n\nThe design is not universally appropriate, and the honest cases against it are:\n\n- **The effect is immediate and the outcome is fast.** A simple parallel A/B test at user level is cheaper and cleaner. Use this design when user-level randomisation is genuinely unavailable.\n- **The rollout cannot be sequenced randomly for good reasons.** If the first cohort must be the design partners who co-built the feature, the order is not random and you are back to [quasi-experimental methods](/docs/quasi-experimental-design-guide).\n- **The window is long relative to market volatility.** A twelve-month rollout through a turbulent period stacks so much secular change against the estimate that the time model does more work than the data supports.\n- **Contamination between clusters is high.** If accounts talk to each other and the intervention is a visible workflow change, control-period accounts may be affected before they cross over.\n\n## A working checklist\n\n1. Confirm the change will reach everyone eventually, and that the rollout will be phased.\n2. Define the cluster: workspace, account, region or segment.\n3. Choose the number of steps and the step length, matched to how fast the outcome moves.\n4. **Randomise which cluster crosses over at which step, and write the assignment down before the first step.**\n5. Pre-commit to modelling period as a term in the analysis.\n6. Instrument every cluster in every period, control periods included.\n7. Field an identical structured study at each step across all clusters.\n8. Analyse the vertical comparisons and the horizontal ones, and report the time trend alongside the effect.\n\n## Frequently asked questions\n\n### What is a stepped wedge cluster randomised trial?\n\nIt is a design in which all clusters begin in the control condition and cross over to the intervention one group at a time, in a randomly assigned order, until every cluster has received it. Each cluster contributes observations both before and after its own crossover, so the design yields within-cluster before-and-after comparisons and contemporaneous between-cluster comparisons at the same time.\n\n### How is this different from a normal phased rollout?\n\nOnly one thing differs, and it is the thing that matters: the order is decided at random and recorded in advance. A conventional phased rollout sequences accounts by size, region, eagerness or account-manager preference, and every one of those criteria is correlated with the outcomes you care about. Randomising the order costs nothing and converts the same rollout into a randomised comparison.\n\n### Is it a quasi-experimental design?\n\nNo, and the distinction is the practical point of the method. Quasi-experimental designs apply when assignment was not random and a counterfactual must be reconstructed from trends or comparison groups. A stepped wedge is genuinely randomised; what is randomised is the timing of exposure rather than who is exposed.\n\n### Do I need a lot of accounts?\n\nFewer than a parallel design requires. Because every cluster acts as its own control before crossover and contributes data in multiple periods, the design is comparatively efficient with small numbers of clusters. Published trials have a median of about 20 clusters, but the logic works with considerably fewer, which is why it suits enterprise portfolios where a parallel split of a dozen accounts would be hopeless.\n\n### What is the main risk?\n\nConfounding with time. Because exposure increases monotonically across the rollout, any other trend over the same window can be mistaken for the effect. The analysis must model the time period explicitly. Reviews of published stepped wedge trials have found that reporting of how time was handled is frequently inadequate, so this is a real failure mode rather than a theoretical one.\n\n### How long should each step be?\n\nLong enough for the outcome to respond and be observed, plus any expected lag before the intervention takes effect. If a workflow change needs six weeks to alter renewal behaviour, a four-week step will attribute movement to the wrong period. Decide the expected lag before the rollout begins and set the step length against it rather than against the release calendar.\n\n## Related Resources\n\n- [Quasi-Experimental Design](/docs/quasi-experimental-design-guide) - what to do when the order genuinely cannot be randomised\n- [The Healthy Adherer Effect](/docs/healthy-adherer-effect-product-research) - the confounding this design prevents by construction\n- [Collider Bias](/docs/collider-bias-product-research) - why measuring only the enthusiastic accounts at the end defeats the design\n- [A/B Testing vs User Research](/docs/ab-testing-vs-user-research) - when to randomise and when to ask\n- [Staged Rollout for AI Features](/docs/ai-staged-rollout-user-research) - risk gating across a release, a different job from causal estimation\n- [Surveillance Bias](/docs/surveillance-bias-detection-research) - why probing depth must be identical across periods\n- [Structured Questions Guide](/docs/structured-questions-guide) - building an instrument that holds still across repeated waves\n\n**Randomise the order of your next rollout.** Koji gives you 10 free interview credits, which is enough to field the first step of a stepped wedge across your control and intervention accounts.","category":"Research Methods","lastModified":"2026-08-14T03:22:39.89955+00:00","metaTitle":"Stepped Wedge Rollout: Randomise When, Not Who (2026)","metaDescription":"You cannot randomise who gets a feature when everyone gets it. You can randomise when. The stepped wedge turns a phased B2B rollout into a randomised trial at almost no extra cost.","keywords":["stepped wedge","cluster randomised trial","randomise rollout order","phased rollout experiment","stepped wedge design","account level experimentation","B2B experimentation"],"aiSummary":"A stepped wedge cluster randomised trial starts every cluster in the control condition and crosses them over to the intervention in a randomly assigned order until all are exposed. It solves the case where a change must reach everyone, so who receives it cannot be randomised but when they receive it can. Originated in the Gambia Hepatitis Intervention Study from 1986, where 17 vaccination teams were randomly assigned starting dates. It suits B2B software because features roll out per account and rollouts are already phased. The main risk is confounding with time, which requires modelling period explicitly; reviews show reporting of time handling is often inadequate.","aiPrerequisites":["Familiarity with A/B testing concepts","Understanding of control groups and confounding"],"aiLearningOutcomes":["Recognise when a phased rollout can be converted into a randomised trial","Lay out a stepped wedge schedule for account-level clusters","Explain why time must be modelled explicitly in the analysis","Judge when a stepped wedge is the wrong choice"],"aiDifficulty":"advanced","aiEstimatedTime":"13 min read"}],"pagination":{"total":1,"returned":1,"offset":0}}