Back to blog
Research

Choice Architecture and Defaults (2026): Why the Variant That Wins the Test Can Lose the Customer

A design change that raises conversion by making the decision easier to make also makes it easier to make badly. The best evidence on nudges at scale says the honest effect is 1.4 points, not 8.7 - and the downstream cost lands in a different quarter, on a different team's dashboard.

Koji

Koji Team

Research · · 12 min read

The previous article in this series argued that a product detail page test measures whether a page closed a sale, not what the shopper came to find out. This one goes further and inverts the sign.

Every conversion team operates on a shared assumption: reduce friction, pre-select the sensible option, and a higher completion rate is straightforwardly good news. That assumption holds for a large class of changes. It fails for one specific class - and that class contains most of the interventions teams are proudest of.

The short answer

A design change that raises conversion by making the decision easier to make also makes it easier to make badly. Conversion and commitment are not the same variable, and a single intervention can move them in opposite directions.

When that happens, your funnel dashboard and your retention dashboard are both correctly reporting the effect of the same change - with opposite signs. The funnel reports it within a week. Retention reports it a quarter later, in a different meeting, owned by a different team, attributed to something else entirely.

To be clear about scope: this is not an article about deceptive design. Manipulative flows are a separate and well-covered problem, and if that is your question, start with dark patterns testing instead. This is about entirely legitimate choice architecture - a sensible default, a shorter form, a pre-selected plan, a simplified comparison - doing something you did not measure.

The evidence base is thinner than the field admits

Before accepting any effect size for a choice-architecture change, it is worth knowing what the best evidence on nudges at scale actually shows. That evidence exists, it is unusually good, and it is sobering.

DellaVigna and Linos, in RCTs to Scale: Comprehensive Evidence from Two Nudge Units (Econometrica, 2022), assembled 126 randomized trials involving 243 nudges and over 23 million participants - every trial run by two of the largest nudge units in the United States. They compared those results against a sample of nudge trials published in academic journals.

The gap is enormous:

  • In the academic journals sample, the average nudge produced an 8.7 percentage point take-up effect - a 33.5% increase over the control mean.
  • In the nudge unit trials, the average effect was 1.4 percentage points, an 8.1% increase.

Roughly a sixfold difference between the literature and practice. The authors then explain it, and the explanation is the part that should change how you read every case study you have ever seen:

  • Controlling for statistical power alone "explains the entire difference" - well-powered academic nudges land at around 1 percentage point, right on top of the practitioner number.
  • Modelling selective publication directly, they estimate that trials with no significant results are written up and published with probability 0.1. Correcting for that pulls the academic average from 8.6 down to 3.2 percentage points.
  • When they asked people to forecast the nudge-unit results, the median prediction was 4 percentage points, nearly triple the truth. Practising nudge professionals were far more accurate, at a median of 1.95 points.

The honest reading is not "nudges do not work." A 1.4 point lift at near-zero marginal cost is an excellent return, and the authors say so. The honest reading is that the effect sizes circulating in conference talks and vendor case studies are drawn from the same selection process that produced the 8.7 figure, and your own uplift will look like the 1.4.

There is one more detail that matters enormously here, and it cuts against my own argument, so it belongs in the open. The authors note that the nudge units "largely rule out default changes, that tend to have larger impacts," which is why they treat 1.4 points as a lower bound. Defaults are the strongest lever in choice architecture - and they are the lever specifically excluded from the best at-scale evidence we have. The intervention with the biggest effect is the one with the least trustworthy evidence base. That is the situation you are making decisions in.

Why the sign flips

A default does two things at once, and teams only measure one of them.

It moves the outcome - more people end up on the pre-selected plan, the larger size, the annual term. That is the measured effect.

It also removes the occasion on which a preference gets formed. A shopper who compares three options and picks one has done work; whatever they end up with, they now hold a reason. A shopper who accepts what was already selected has done no work and holds no reason. Both are counted identically in the conversion numerator. They behave completely differently afterwards.

That is the whole mechanism, and it produces a clean prediction: interventions that raise conversion by answering a question should improve downstream outcomes, while interventions that raise conversion by removing the question should degrade them.

Type of changeWhat it does to the decisionConversionLikely downstream effect
Adding the missing spec, size chart, or return policyResolves an open questionUpBetter - fewer mismatched purchases
Removing a genuinely redundant stepNo effect on the decisionUpNeutral
Pre-selecting the option most people chooseRemoves the comparisonUpWorse for the minority who needed to compare
Simplifying a comparison by cutting attributesRemoves the trade-offUpWorse where the cut attribute was the deciding one
Deferring a commitment disclosure to laterPostpones the decisionUpWorst - the decision surfaces as a cancellation

Rows one and two are why friction reduction has such a good reputation; they are real, and they are free. Rows three through five are where the reputation gets borrowed to cover something else.

The downstream cost is measurable, and in ecommerce it has a price tag. The National Retail Federation's 2025 Retail Returns Landscape, based on a summer 2025 survey of 2,006 consumers and 358 ecommerce professionals at US merchants above $500 million in revenue, put 2025 returns at $849.9 billion on a 15.8% return rate, with 19.3% of online sales expected to be returned. And the penalty compounds: 71% of consumers said they are less likely to shop with a retailer again after a poor returns experience, up from 67% in 2024.

A purchase completed without deliberation is a strong candidate to become one of those returns. Nothing in the A/B test result will tell you which ones.

The measurement gap nobody closes

The reason this persists is not stupidity. It is a mismatch of clocks and owners.

An A/B test reaches significance in days or weeks, on a binary outcome, owned by growth. The consequence appears in returns, refunds, cancellations, or support tickets over the following one to two quarters, owned by operations, retention, or CX. By the time the second signal arrives, the test has been shipped, written up, and cited in three other proposals. Nobody re-opens it, because nobody has a reason to connect the two.

Worse, the second signal is usually not attributable even if somebody tries. Returns and cancellations are rarely tagged with the acquisition variant. The counterfactual is gone.

So the failure mode is not that teams get the wrong answer. It is that the question is never asked, by anyone, at any point.

How to research it

You are trying to establish whether people who converted under the new design hold a reason for what they chose. That is not observable in behavioural data and it is not reliably answerable on a form, because "did you consider the alternatives?" is a question almost everybody answers yes to.

The move that works is the re-decision probe. Do not ask whether they compared. Ask them to reconstruct what they compared against, and check whether the reconstruction has any content in it.

Somebody who deliberated names the rejected option and gives a reason. Somebody who accepted a default says something like "it seemed like the standard one" and cannot go further under follow-up. The distinction is completely invisible to a checkbox and completely obvious to a conversation - which is exactly why this needs an AI moderator that probes rather than a static form that records.

Run it on both arms of a live test, and the structure carries the argument:

What you needQuestion typeExample
Whether a comparison occurredyes_noDid you look at any option other than the one you picked?
What it was compared againstopen_endedWhich other option did you consider, and why did you rule it out?
Strength of the held reasonscaleHow confident are you that this was the right choice for you?
Whether the default was noticedsingle_choiceWas the option you chose already selected when you arrived?
What would have changed the pickmultiple_choiceWhich of these would have made you choose differently?
The real decision driversrankingRank price, commitment length, flexibility, features, and speed

The open_ended item is the whole instrument. Everything else is scaffolding that makes the result countable.

The output is a number you can put next to your conversion lift: the share of converters in each arm who can name a rejected alternative. If the winning variant converts 6% better and carries 20 points fewer reasoned choices, you have not found an uplift. You have found a trade, and now you can price it.

Where Koji fits

This study is impossible with legacy tooling on any sensible timeline. Typeform and SurveyMonkey cannot follow up on a vague answer, so everybody passes the comparison question. UserTesting and dscout can get depth, but scheduling moderated sessions against a live experiment means your results arrive after the decision to ship. Dovetail organises transcripts you still have to pay somebody to produce.

Koji runs AI-moderated voice interviews against both arms of a running test, at the volume you need, in the window you actually have. The moderator probes every vague answer the same way, so there is no interviewer drift between arm A and arm B - which matters more here than in almost any other study, because the entire finding is a comparison between two groups. Thematic analysis runs automatically as transcripts land, and a one-click report gives you the reasoned-choice rate per arm without anybody coding transcripts by hand.

Pair it with structured questions for the countable side, conjoint analysis when you need to model the trade-offs directly, and cancel-flow exit interviews to catch the downstream half of the effect where it lands.

What to do in the next quarter

  • Pick your three most celebrated conversion wins from the last year. For each, ask which of the five rows above it belongs to.
  • For anything in rows three to five, tag the cohort and check returns, refunds, or cancellations against the control cohort. If nobody tagged it, that is the finding.
  • Add the reasoned-choice rate to the standard readout of any test that changes what is pre-selected or what is removed.
  • Discount every external uplift figure you are quoted. The best evidence we have says the honest number is closer to 1.4 points than to 8.7.

The variant that wins the test is not always the variant you want to ship. You just have to measure long enough to tell.

Frequently Asked Questions

Is this saying defaults are bad?

No. Defaults are one of the most effective and cheapest design tools available, and for the majority of users a well-chosen default is genuinely helpful. The argument is narrower: a default converts two distinct groups - people it helped and people it prevented from thinking - and your conversion metric adds them together. You need to know the ratio before you decide whether the win is real.

How is this different from dark patterns testing?

Dark patterns testing asks whether a flow is deceptive and whether it would survive regulatory scrutiny. This asks whether an entirely honest design has an unmeasured downstream cost. A default can be fully disclosed, legally impeccable, and still produce customers who cannot say why they bought. The two studies use similar methods and answer different questions.

What is the reasoned-choice rate?

It is the share of people who converted who can name a specific alternative they rejected and give a reason for rejecting it. Measure it separately in each arm of a test. A variant that lifts conversion while lowering the reasoned-choice rate has traded future commitment for present completion, and you can then decide whether that trade is worth making.

Why can't we just look at the retention data?

Usually because the cohort was never tagged. Returns, refunds, and cancellations are recorded against the customer, not against the experiment variant that acquired them, and by the time the downstream signal appears the test has been concluded and the assignment data has often been discarded. Tagging the cohort at test time costs almost nothing and is the single highest-leverage change here.

How large a sample do we need for the re-decision probe?

Because the outcome is a proportion compared across two arms, plan for 60 to 100 completed conversations per arm to detect a difference of 15 points or more with reasonable confidence. That is impractical with moderated interviews on an experiment timeline, which is why AI moderation is what makes this study feasible at all rather than merely nicer.

Does this apply to B2B products as well as ecommerce?

Yes, and often more strongly. In B2B the pre-selected plan, the default seat count, and the default contract term shape a decision that gets revisited at renewal by somebody who may not have made it. The downstream signal is slower - a renewal cycle rather than a return window - which makes the attribution problem worse, not better.

Run your first AI-moderated study in 10 minutes

10 free credits on signup. No credit card required.

GDPR compliantEU or US data residencyNo AI training on your data
Koji

Koji Team

Research

Share this article

Keep reading