{"site":{"name":"Koji","description":"AI-native customer research platform that helps teams conduct, analyze, and synthesize customer interviews at scale.","url":"https://www.koji.so","contentTypes":["blog","documentation"],"lastUpdated":"2026-08-19T09:55:52.517Z"},"content":[{"type":"blog","id":"a9bb0994-5d74-448f-a29c-b0cf1ad6e633","slug":"choice-architecture-defaults-research-2026","title":"Choice Architecture and Defaults (2026): Why the Variant That Wins the Test Can Lose the Customer","url":"https://www.koji.so/blog/choice-architecture-defaults-research-2026","summary":"Choice architecture changes that raise conversion by removing a comparison can lower downstream commitment, so the same intervention shows up with opposite signs on the funnel and retention dashboards a quarter apart. DellaVigna and Linos' Econometrica study of 126 trials, 243 nudges and 23 million participants found academic nudges average 8.7 percentage points against 1.4 points for nudge units, with statistical power and publication bias explaining the entire gap - and defaults, the strongest lever, are largely excluded from that at-scale evidence. The proposed measure is the reasoned-choice rate: the share of converters in each arm who can name a rejected alternative.","content":"The previous article in this series argued that a [product detail page](/blog/product-detail-page-research-2026) test measures whether a page closed a sale, not what the shopper came to find out. This one goes further and inverts the sign.\n\nEvery conversion team operates on a shared assumption: reduce friction, pre-select the sensible option, and a higher completion rate is straightforwardly good news. That assumption holds for a large class of changes. It fails for one specific class - and that class contains most of the interventions teams are proudest of.\n\n## The short answer\n\n**A design change that raises conversion by making the decision easier to make also makes it easier to make badly.** Conversion and commitment are not the same variable, and a single intervention can move them in opposite directions.\n\nWhen that happens, your funnel dashboard and your retention dashboard are both correctly reporting the effect of the same change - with opposite signs. The funnel reports it within a week. Retention reports it a quarter later, in a different meeting, owned by a different team, attributed to something else entirely.\n\nTo be clear about scope: this is not an article about deceptive design. Manipulative flows are a separate and well-covered problem, and if that is your question, start with [dark patterns testing](/docs/dark-patterns-user-testing) instead. **This is about entirely legitimate choice architecture** - a sensible default, a shorter form, a pre-selected plan, a simplified comparison - doing something you did not measure.\n\n## The evidence base is thinner than the field admits\n\nBefore accepting any effect size for a choice-architecture change, it is worth knowing what the best evidence on nudges at scale actually shows. That evidence exists, it is unusually good, and it is sobering.\n\nDellaVigna and Linos, in *RCTs to Scale: Comprehensive Evidence from Two Nudge Units* (Econometrica, 2022), assembled **126 randomized trials involving 243 nudges and over 23 million participants** - every trial run by two of the largest nudge units in the United States. They compared those results against a sample of nudge trials published in academic journals.\n\nThe gap is enormous:\n\n- In the **academic journals** sample, the average nudge produced an **8.7 percentage point** take-up effect - a **33.5% increase** over the control mean.\n- In the **nudge unit** trials, the average effect was **1.4 percentage points**, an **8.1% increase**.\n\nRoughly a sixfold difference between the literature and practice. The authors then explain it, and the explanation is the part that should change how you read every case study you have ever seen:\n\n- Controlling for statistical power alone **\"explains the entire difference\"** - well-powered academic nudges land at **around 1 percentage point**, right on top of the practitioner number.\n- Modelling selective publication directly, they estimate that trials with no significant results are written up and published **with probability 0.1**. Correcting for that pulls the academic average from **8.6 down to 3.2 percentage points**.\n- When they asked people to forecast the nudge-unit results, the median prediction was **4 percentage points**, nearly triple the truth. Practising nudge professionals were far more accurate, at a median of **1.95 points**.\n\nThe honest reading is not \"nudges do not work.\" A 1.4 point lift at near-zero marginal cost is an excellent return, and the authors say so. The honest reading is that **the effect sizes circulating in conference talks and vendor case studies are drawn from the same selection process that produced the 8.7 figure**, and your own uplift will look like the 1.4.\n\nThere is one more detail that matters enormously here, and it cuts against my own argument, so it belongs in the open. The authors note that the nudge units **\"largely rule out default changes, that tend to have larger impacts,\"** which is why they treat 1.4 points as a lower bound. Defaults are the strongest lever in choice architecture - and they are the lever specifically excluded from the best at-scale evidence we have. **The intervention with the biggest effect is the one with the least trustworthy evidence base.** That is the situation you are making decisions in.\n\n## Why the sign flips\n\nA default does two things at once, and teams only measure one of them.\n\nIt moves the outcome - more people end up on the pre-selected plan, the larger size, the annual term. That is the measured effect.\n\nIt also **removes the occasion on which a preference gets formed**. A shopper who compares three options and picks one has done work; whatever they end up with, they now hold a reason. A shopper who accepts what was already selected has done no work and holds no reason. Both are counted identically in the conversion numerator. They behave completely differently afterwards.\n\nThat is the whole mechanism, and it produces a clean prediction: interventions that raise conversion by *answering* a question should improve downstream outcomes, while interventions that raise conversion by *removing* the question should degrade them.\n\n| Type of change | What it does to the decision | Conversion | Likely downstream effect |\n| --- | --- | --- | --- |\n| Adding the missing spec, size chart, or return policy | Resolves an open question | Up | Better - fewer mismatched purchases |\n| Removing a genuinely redundant step | No effect on the decision | Up | Neutral |\n| Pre-selecting the option most people choose | Removes the comparison | Up | Worse for the minority who needed to compare |\n| Simplifying a comparison by cutting attributes | Removes the trade-off | Up | Worse where the cut attribute was the deciding one |\n| Deferring a commitment disclosure to later | Postpones the decision | Up | Worst - the decision surfaces as a cancellation |\n\nRows one and two are why friction reduction has such a good reputation; they are real, and they are free. Rows three through five are where the reputation gets borrowed to cover something else.\n\nThe downstream cost is measurable, and in ecommerce it has a price tag. The National Retail Federation's **2025 Retail Returns Landscape**, based on a summer 2025 survey of **2,006 consumers** and **358 ecommerce professionals** at US merchants above $500 million in revenue, put 2025 returns at **$849.9 billion** on a **15.8%** return rate, with **19.3% of online sales** expected to be returned. And the penalty compounds: **71%** of consumers said they are less likely to shop with a retailer again after a poor returns experience, up from **67% in 2024**.\n\nA purchase completed without deliberation is a strong candidate to become one of those returns. Nothing in the A/B test result will tell you which ones.\n\n## The measurement gap nobody closes\n\nThe reason this persists is not stupidity. It is a mismatch of clocks and owners.\n\nAn A/B test reaches significance in days or weeks, on a binary outcome, owned by growth. The consequence appears in returns, refunds, cancellations, or support tickets over the following one to two quarters, owned by operations, retention, or CX. By the time the second signal arrives, the test has been shipped, written up, and cited in three other proposals. Nobody re-opens it, because nobody has a reason to connect the two.\n\nWorse, the second signal is usually not attributable even if somebody tries. Returns and cancellations are rarely tagged with the acquisition variant. The counterfactual is gone.\n\n**So the failure mode is not that teams get the wrong answer. It is that the question is never asked, by anyone, at any point.**\n\n## How to research it\n\nYou are trying to establish whether people who converted under the new design **hold a reason** for what they chose. That is not observable in behavioural data and it is not reliably answerable on a form, because \"did you consider the alternatives?\" is a question almost everybody answers yes to.\n\nThe move that works is the **re-decision probe**. Do not ask whether they compared. Ask them to reconstruct what they compared *against*, and check whether the reconstruction has any content in it.\n\nSomebody who deliberated names the rejected option and gives a reason. Somebody who accepted a default says something like \"it seemed like the standard one\" and cannot go further under follow-up. The distinction is completely invisible to a checkbox and completely obvious to a conversation - which is exactly why this needs an AI moderator that probes rather than a static form that records.\n\nRun it on both arms of a live test, and the structure carries the argument:\n\n| What you need | Question type | Example |\n| --- | --- | --- |\n| Whether a comparison occurred | `yes_no` | Did you look at any option other than the one you picked? |\n| What it was compared against | `open_ended` | Which other option did you consider, and why did you rule it out? |\n| Strength of the held reason | `scale` | How confident are you that this was the right choice for you? |\n| Whether the default was noticed | `single_choice` | Was the option you chose already selected when you arrived? |\n| What would have changed the pick | `multiple_choice` | Which of these would have made you choose differently? |\n| The real decision drivers | `ranking` | Rank price, commitment length, flexibility, features, and speed |\n\nThe `open_ended` item is the whole instrument. Everything else is scaffolding that makes the result countable.\n\nThe output is a number you can put next to your conversion lift: **the share of converters in each arm who can name a rejected alternative.** If the winning variant converts 6% better and carries 20 points fewer reasoned choices, you have not found an uplift. You have found a trade, and now you can price it.\n\n## Where Koji fits\n\nThis study is impossible with legacy tooling on any sensible timeline. Typeform and SurveyMonkey cannot follow up on a vague answer, so everybody passes the comparison question. UserTesting and dscout can get depth, but scheduling moderated sessions against a live experiment means your results arrive after the decision to ship. Dovetail organises transcripts you still have to pay somebody to produce.\n\nKoji runs AI-moderated voice interviews against both arms of a running test, at the volume you need, in the window you actually have. The moderator probes every vague answer the same way, so there is no interviewer drift between arm A and arm B - which matters more here than in almost any other study, because the entire finding is a comparison between two groups. Thematic analysis runs automatically as transcripts land, and a one-click report gives you the reasoned-choice rate per arm without anybody coding transcripts by hand.\n\nPair it with [structured questions](/docs/structured-questions-guide) for the countable side, [conjoint analysis](/docs/conjoint-analysis-guide) when you need to model the trade-offs directly, and [cancel-flow exit interviews](/docs/cancel-flow-exit-interview) to catch the downstream half of the effect where it lands.\n\n## What to do in the next quarter\n\n- Pick your three most celebrated conversion wins from the last year. For each, ask which of the five rows above it belongs to.\n- For anything in rows three to five, tag the cohort and check returns, refunds, or cancellations against the control cohort. If nobody tagged it, that is the finding.\n- Add the reasoned-choice rate to the standard readout of any test that changes what is pre-selected or what is removed.\n- Discount every external uplift figure you are quoted. The best evidence we have says the honest number is closer to 1.4 points than to 8.7.\n\nThe variant that wins the test is not always the variant you want to ship. You just have to measure long enough to tell.\n\n## Related reading\n\n- [Product Detail Page Research (2026)](/blog/product-detail-page-research-2026)\n- [Subscribe and Save (2026)](/blog/replenishment-subscription-research-2026)\n- [Share of Search and the Digital Shelf (2026)](/blog/share-of-search-digital-shelf-2026)\n- [Personalized Pricing in 2026](/blog/personalized-pricing-research-2026)\n- [Customer Research for Ecommerce Brands](/blog/customer-research-for-ecommerce-2026)\n- [How to Reduce Return Rate (2026)](/blog/return-rate-reduction-research-2026) - why the same metric falls for opposite reasons.\n\n## Frequently Asked Questions\n\n### Is this saying defaults are bad?\n\nNo. Defaults are one of the most effective and cheapest design tools available, and for the majority of users a well-chosen default is genuinely helpful. The argument is narrower: a default converts two distinct groups - people it helped and people it prevented from thinking - and your conversion metric adds them together. You need to know the ratio before you decide whether the win is real.\n\n### How is this different from dark patterns testing?\n\nDark patterns testing asks whether a flow is deceptive and whether it would survive regulatory scrutiny. This asks whether an entirely honest design has an unmeasured downstream cost. A default can be fully disclosed, legally impeccable, and still produce customers who cannot say why they bought. The two studies use similar methods and answer different questions.\n\n### What is the reasoned-choice rate?\n\nIt is the share of people who converted who can name a specific alternative they rejected and give a reason for rejecting it. Measure it separately in each arm of a test. A variant that lifts conversion while lowering the reasoned-choice rate has traded future commitment for present completion, and you can then decide whether that trade is worth making.\n\n### Why can't we just look at the retention data?\n\nUsually because the cohort was never tagged. Returns, refunds, and cancellations are recorded against the customer, not against the experiment variant that acquired them, and by the time the downstream signal appears the test has been concluded and the assignment data has often been discarded. Tagging the cohort at test time costs almost nothing and is the single highest-leverage change here.\n\n### How large a sample do we need for the re-decision probe?\n\nBecause the outcome is a proportion compared across two arms, plan for 60 to 100 completed conversations per arm to detect a difference of 15 points or more with reasonable confidence. That is impractical with moderated interviews on an experiment timeline, which is why AI moderation is what makes this study feasible at all rather than merely nicer.\n\n### Does this apply to B2B products as well as ecommerce?\n\nYes, and often more strongly. In B2B the pre-selected plan, the default seat count, and the default contract term shape a decision that gets revisited at renewal by somebody who may not have made it. The downstream signal is slower - a renewal cycle rather than a return window - which makes the attribution problem worse, not better.\n","category":"Research","lastModified":"2026-08-18T03:23:58.984331+00:00","metaTitle":"Choice Architecture and Defaults (2026): The Conversion Win That Costs You Later","metaDescription":"Conversion and commitment can move in opposite directions. What DellaVigna and Linos' 126-trial nudge study says about real effect sizes, and how to measure the reasoned-choice rate.","keywords":["choice architecture research","default option research","nudge effect size","consumer defaults","conversion rate optimization research","deliberation and conversion"],"aiSummary":"Choice architecture changes that raise conversion by removing a comparison can lower downstream commitment, so the same intervention shows up with opposite signs on the funnel and retention dashboards a quarter apart. DellaVigna and Linos' Econometrica study of 126 trials, 243 nudges and 23 million participants found academic nudges average 8.7 percentage points against 1.4 points for nudge units, with statistical power and publication bias explaining the entire gap - and defaults, the strongest lever, are largely excluded from that at-scale evidence. The proposed measure is the reasoned-choice rate: the share of converters in each arm who can name a rejected alternative.","aiKeywords":["choice architecture","default effect","nudge units","DellaVigna Linos","publication bias","reasoned choice rate","conversion testing","AI moderated interviews"],"aiContentType":"guide","faqItems":[{"answer":"No. Defaults are one of the most effective and cheapest design tools available, and for the majority of users a well-chosen default is genuinely helpful. The argument is narrower: a default converts two distinct groups - people it helped and people it prevented from thinking - and your conversion metric adds them together. You need to know the ratio before you decide whether the win is real.","question":"Is this saying defaults are bad?"},{"answer":"Dark patterns testing asks whether a flow is deceptive and whether it would survive regulatory scrutiny. This asks whether an entirely honest design has an unmeasured downstream cost. A default can be fully disclosed, legally impeccable, and still produce customers who cannot say why they bought. The two studies use similar methods and answer different questions.","question":"How is this different from dark patterns testing?"},{"answer":"It is the share of people who converted who can name a specific alternative they rejected and give a reason for rejecting it. Measure it separately in each arm of a test. A variant that lifts conversion while lowering the reasoned-choice rate has traded future commitment for present completion, and you can then decide whether that trade is worth making.","question":"What is the reasoned-choice rate?"},{"answer":"Usually because the cohort was never tagged. Returns, refunds, and cancellations are recorded against the customer, not against the experiment variant that acquired them, and by the time the downstream signal appears the test has been concluded and the assignment data has often been discarded. Tagging the cohort at test time costs almost nothing and is the single highest-leverage change here.","question":"Why can't we just look at the retention data?"},{"answer":"Because the outcome is a proportion compared across two arms, plan for 60 to 100 completed conversations per arm to detect a difference of 15 points or more with reasonable confidence. That is impractical with moderated interviews on an experiment timeline, which is why AI moderation is what makes this study feasible at all rather than merely nicer.","question":"How large a sample do we need for the re-decision probe?"},{"answer":"Yes, and often more strongly. In B2B the pre-selected plan, the default seat count, and the default contract term shape a decision that gets revisited at renewal by somebody who may not have made it. The downstream signal is slower - a renewal cycle rather than a return window - which makes the attribution problem worse, not better.","question":"Does this apply to B2B products as well as ecommerce?"}],"relatedTopics":["behavioral science","conversion research","experimentation","product returns","ecommerce research"]}],"pagination":{"total":1,"returned":1,"offset":0}}