Back to docs
Research Methods

The Stepped Wedge: How to Randomise a Rollout You Cannot Randomise (2026)

You cannot randomise who gets the feature, because everyone is getting it. You can randomise when. The stepped wedge turns a phased rollout into a randomised trial at almost no extra cost, and it is the design B2B teams are already accidentally halfway to running.

The standard advice for a change that ships to everyone is to give up on randomisation and reconstruct a counterfactual afterwards. That advice is often wrong, because it answers the wrong question. You cannot randomise WHO receives the change. You can almost always randomise WHEN they receive it, and that is enough to get a randomised comparison.

This is the stepped wedge cluster randomised trial. Every unit starts in the control condition, units cross over to the intervention in a randomly assigned order, and by the end everyone has it. Nobody is denied anything. The rollout you were going to do anyway becomes the experiment.

If you have already read quasi-experimental design, this is the design that rescues a case that article correctly concedes. Quasi-experimental methods construct a counterfactual from pre-existing trends or untreated comparison groups because assignment was not random. The stepped wedge does not need to construct one, because the assignment was random after all. The difference between the two is not statistical sophistication. It is whether somebody thought about the rollout order before the rollout started.

The design

Hemming and colleagues describe it in the BMJ (2015;350:h391) as involving "random and sequential crossover of clusters from control to intervention until all clusters are exposed". Every cluster provides both before and after observations, and every cluster switches, but not at the same time.

A four-step rollout across twelve accounts looks like this, where C is control and I is intervention:

GroupPeriod 1Period 2Period 3Period 4Period 5
Accounts 1-3CIIII
Accounts 4-6CCIII
Accounts 7-9CCCII
Accounts 10-12CCCCI

The group assignment to rows is the randomised part. Reading down any column gives you a contemporaneous comparison between accounts that have the feature and accounts that do not. Reading across any row gives you a within-account before-and-after. The design supplies both kinds of comparison from a rollout that most teams would have sequenced by account size or by whoever asked loudest.

Where it came from

The first stepped wedge trial was the Gambia Hepatitis Intervention Study, launched in July 1986. The goal was to evaluate adding hepatitis B vaccine to the existing childhood immunisation schedule. Instantaneous nationwide vaccination was impossible for logistical and financial reasons, so the 17 teams delivering the existing vaccines were randomly assigned starting dates for adding the new one, phased over about four years.

The reasoning translates to product work almost word for word. The rollout had to be phased because of capacity. Since it had to be phased, the order had to be decided somehow. Deciding it at random cost nothing and produced an unvaccinated comparison group at every point in time, purely as a by-product.

Why this fits B2B software better than it fits medicine

Stepped wedge designs are used in health services research because interventions there are often delivered to whole hospitals or clinics rather than to individuals. That is exactly the shape of B2B software.

  • Features are frequently enabled per workspace, per account or per region, not per user.
  • Rollouts are already phased, because support, training and infrastructure capacity are finite.
  • Everyone is contractually or politically entitled to the feature eventually, so withholding it from a control group indefinitely is not an option.
  • The unit that experiences the change is the team, and outcomes such as retention and expansion are measured at account level anyway.

Every one of those facts is usually cited as a reason experimentation is impossible. They are the exact conditions the stepped wedge was invented for.

Objection to experimentingWhat the stepped wedge does with it
"We cannot withhold it from customers"Nothing is withheld. Everyone receives it, in a random order
"It ships to the whole account, not to users"The account is the cluster. That is the intended unit
"We only have 15 enterprise accounts"Each account contributes both control and intervention periods, so the design is unusually efficient at small cluster counts
"The rollout is phased for support capacity"Then an order already exists. Randomising it is free
"Leadership will not wait for a trial"The trial finishes when the rollout finishes. It adds no delay

The small-sample point deserves emphasis because it is the most common reason B2B teams believe experimentation is closed to them. In a parallel design, 15 accounts split into two groups is a hopeless comparison. In a stepped wedge, each account serves as its own control before it crosses over, and the design accumulates comparisons at every step.

The cost: time is now a confounder

The stepped wedge is not free, and the price is specific. Because the intervention is perfectly correlated with time, anything else that changes over the rollout window is confounded with the feature. A seasonal uplift, a pricing change, a competitor exit or a support reorganisation will all masquerade as the effect if the analysis ignores time.

This is why the design requires modelling the underlying temporal trend rather than simply comparing before against after. Hemming and Lilford note that results can "transpire to be the result of a positive underlying temporal trend", and that a large number of published studies do not report how they allowed for time in the design or the analysis.

The reporting record is genuinely poor. In a review of 123 studies using the design, of which 39 were completed trial reports, Grayling, Wason and Mander found that reporting quality on individual criteria fell as low as 15.4 percent, with a median of 66.7 percent (Trials 2017;18:33). Only 25 of the 39 completed trials, or 64.1 percent, gave a rationale for choosing the stepped wedge at all. The median trial used 20.5 clusters and 9 steps. The CONSORT reporting extension for these trials (Hemming, Taljaard and colleagues, BMJ 2018;363:k1614) exists partly because of this, and it requires investigators to justify the design choice explicitly.

The lesson for a product team is to decide three things in writing before the first account crosses over:

  1. How many steps, and how long is each? A step must be long enough for the outcome to be observable within it.
  2. How will time be modelled? At minimum, period should enter the analysis as a term. Without it the estimate is not interpretable.
  3. What is the lag before the effect is expected? If a feature takes six weeks to change behaviour, a four-week step attributes the change to the wrong period.

What to measure at each step, and why teams abandon this

The design's weak point in practice is not the statistics. It is that a stepped wedge needs measurement from every cluster in every period, including the periods where nothing has happened yet. Those control-period observations are what the whole design rests on.

Behavioural telemetry covers part of it automatically. The part teams give up on is the human evidence, because interviewing a sample of accounts at five separate points across a rollout is, with traditional methods, five recruitment cycles, five scheduling rounds and five analysis passes. Most teams run interviews once, at the end, with the accounts that are happiest to talk, which reintroduces exactly the conditioning problem the design was supposed to avoid.

This is where the economics have genuinely changed. Koji makes the repeated qualitative arm affordable:

  • AI-moderated interviews can run at every step across every cluster, including the accounts still in the control condition, without consuming a researcher week per wave.
  • Identical moderation across periods and clusters. This is not a convenience point. If wave 5 is probed more thoroughly than wave 1 because the moderator has learned what is interesting, the apparent effect includes the moderator's learning curve. See surveillance bias for how much damage unequal probing does.
  • Voice interviews reach operational users inside customer accounts who will never join a scheduled call.
  • Automatic thematic analysis makes it feasible to compare open-ended themes across ten cells rather than reading them once at the end.
  • Real-time reporting means a step can inform the next step, which is the entire point of a phased rollout.

The six structured question types carry unusual weight in this design. Koji supports open_ended, scale, single_choice, multiple_choice, ranking and yes_no. Because a stepped wedge compares across time, you need items whose meaning does not drift between period 1 and period 5, and closed types are the ones that hold still. A fixed scale item on workflow confidence, a yes_no on whether a workaround is still in use, a single_choice on the primary tool for the task, and a ranking of priorities, all repeated identically at every step, produce a clean time series alongside the telemetry. The open_ended responses then explain the movements rather than having to detect them. The structured questions guide covers building an instrument that survives repetition.

When not to use it

The design is not universally appropriate, and the honest cases against it are:

  • The effect is immediate and the outcome is fast. A simple parallel A/B test at user level is cheaper and cleaner. Use this design when user-level randomisation is genuinely unavailable.
  • The rollout cannot be sequenced randomly for good reasons. If the first cohort must be the design partners who co-built the feature, the order is not random and you are back to quasi-experimental methods.
  • The window is long relative to market volatility. A twelve-month rollout through a turbulent period stacks so much secular change against the estimate that the time model does more work than the data supports.
  • Contamination between clusters is high. If accounts talk to each other and the intervention is a visible workflow change, control-period accounts may be affected before they cross over.

A working checklist

  1. Confirm the change will reach everyone eventually, and that the rollout will be phased.
  2. Define the cluster: workspace, account, region or segment.
  3. Choose the number of steps and the step length, matched to how fast the outcome moves.
  4. Randomise which cluster crosses over at which step, and write the assignment down before the first step.
  5. Pre-commit to modelling period as a term in the analysis.
  6. Instrument every cluster in every period, control periods included.
  7. Field an identical structured study at each step across all clusters.
  8. Analyse the vertical comparisons and the horizontal ones, and report the time trend alongside the effect.

Frequently asked questions

What is a stepped wedge cluster randomised trial?

It is a design in which all clusters begin in the control condition and cross over to the intervention one group at a time, in a randomly assigned order, until every cluster has received it. Each cluster contributes observations both before and after its own crossover, so the design yields within-cluster before-and-after comparisons and contemporaneous between-cluster comparisons at the same time.

How is this different from a normal phased rollout?

Only one thing differs, and it is the thing that matters: the order is decided at random and recorded in advance. A conventional phased rollout sequences accounts by size, region, eagerness or account-manager preference, and every one of those criteria is correlated with the outcomes you care about. Randomising the order costs nothing and converts the same rollout into a randomised comparison.

Is it a quasi-experimental design?

No, and the distinction is the practical point of the method. Quasi-experimental designs apply when assignment was not random and a counterfactual must be reconstructed from trends or comparison groups. A stepped wedge is genuinely randomised; what is randomised is the timing of exposure rather than who is exposed.

Do I need a lot of accounts?

Fewer than a parallel design requires. Because every cluster acts as its own control before crossover and contributes data in multiple periods, the design is comparatively efficient with small numbers of clusters. Published trials have a median of about 20 clusters, but the logic works with considerably fewer, which is why it suits enterprise portfolios where a parallel split of a dozen accounts would be hopeless.

What is the main risk?

Confounding with time. Because exposure increases monotonically across the rollout, any other trend over the same window can be mistaken for the effect. The analysis must model the time period explicitly. Reviews of published stepped wedge trials have found that reporting of how time was handled is frequently inadequate, so this is a real failure mode rather than a theoretical one.

How long should each step be?

Long enough for the outcome to respond and be observed, plus any expected lag before the intervention takes effect. If a workflow change needs six weeks to alter renewal behaviour, a four-week step will attribute movement to the wrong period. Decide the expected lag before the rollout begins and set the step length against it rather than against the release calendar.

Related Resources

Randomise the order of your next rollout. Koji gives you 10 free interview credits, which is enough to field the first step of a stepped wedge across your control and intervention accounts.

Related Articles

A/B Testing vs. User Research: When to Use Each (And When to Use Both)

Understand when A/B testing and qualitative user research each shine, and how to combine them for better product decisions. Includes framework for choosing methods, real case studies, and how AI interviews make mixed methods accessible.

Staged Rollout for AI Features: Shadow Mode, Canary, and Kill Switches (2026)

A research-first guide to staging an AI feature launch. What shadow mode can and cannot measure, what to ask users at each canary ring, how to pre-register rollback thresholds, and why the EU AI Act made the kill switch a legal requirement.

Collider Bias: When Adding a Control Variable Creates the Correlation (2026)

Most research advice tells you to control for more variables. Collider bias is the case where controlling, filtering or segmenting manufactures an association that does not exist. Here is how to recognise it before it reaches a roadmap.

The Healthy Adherer Effect: Why Users Who Finish Onboarding Always Retain Better (2026)

Users who complete your onboarding checklist retain better. So do users who adhere to a placebo. The healthy adherer effect explains why adoption metrics overstate feature impact, why adjusting for covariates does not fix it, and what to do instead.

Immortal Time Bias: Why Feature Adopters Always Look More Loyal Than They Are (2026)

Immortal time bias makes every feature-adoption retention chart overstate the feature. Learn how the bias works, why product data is the worst case, and the three fixes.

Quasi-Experimental Design: How to Measure Impact When You Cannot Run an A/B Test (2026)

Most product decisions cannot be randomised. Quasi-experimental designs give you a defensible causal answer anyway. Learn which of the three designs your situation calls for, how to write the impact model before the data arrives, and why interviews are the cheapest confounder detector you have.

Structured Questions in AI Interviews

Mix quantitative data collection — scales, ratings, multiple choice, ranking — with AI-powered conversational follow-up in a single interview.

Surveillance Bias: Why the Team That Measures Best Looks Worst (2026)

The harder you look, the more you find. Surveillance and lead time bias make well-instrumented teams look worse and useless interventions look effective. Here is how to tell the difference.