A return rate that goes down is not automatically good news. It falls when your listings get more accurate, and it falls when returning becomes annoying enough that people give up. Those are opposite outcomes with identical dashboards. One of them is a customer experience win. The other is a customer loss programme running under a green arrow, and the metric is structurally incapable of telling you which one you just shipped.
This is not hypothetical. It is what the US retail industry did to itself in 2025, and the numbers are public.
The short answer
The return rate is a ratio with at least five different ways of falling, and only two of them are good. To reduce returns safely you have to pair the return rate with a second measurement that moves in opposite directions under accuracy and under suppression. The cleanest pairing is the repeat purchase rate of the customers who kept the item. Without it, a returns reduction programme is unfalsifiable.
The year the industry ran this experiment
Two figures from the National Retail Federation and Happy Returns 2025 Retail Returns Landscape, published 15 October 2025 and based on 2,006 consumers plus 358 ecommerce professionals at US merchants above $500 million in revenue, surveyed in summer 2025:
- The US return rate fell from 16.9% of sales in 2024 to 15.8% in 2025, with the total dropping from $890 billion to $849.9 billion.
- Over the same period, the share of US merchants charging for at least some return options rose from 66% to 72%.
Those two facts sit in the same report. The return rate improved in the same year that six percent of the market started charging for returns. The reasons merchants gave for the fees were increases in the cost of processing returns (40%), increases in carrier shipping costs (40%) and economic uncertainty and tariff risk (33%). Note what is absent from that list: nobody said they introduced fees because their listings had become more accurate.
Now put a third number beside them. In the same survey, 71% of consumers said they are less likely to shop with a retailer again after a poor returns experience, up from 67% in 2024, and 80% say they share a negative returns experience. Meanwhile 82% say free returns are an important purchase consideration, up from 76%.
So the demonstrated sequence is: returns got more expensive and more annoying for customers, the return rate improved, and the customer's stated sensitivity to exactly that friction went up in the same year. A returns dashboard would have recorded that entire year as a success.
Why the sign flips
The return rate is returns / sales. Here is the full set of ways it goes down.
| How the rate falls | Mechanism | What it means for the business |
|---|---|---|
| Listings answer more questions before purchase | Fewer mismatches created | Genuinely good. Fewer returns and better purchases |
| Longer return window | Endowment effect, attachment grows | Good, and counter-intuitive. See below |
| Return fees and added friction | Return suppressed, not prevented | Bad. You kept the revenue and lost the customer |
| Shorter window or narrowed eligibility | Return refused, not prevented | Bad, and adds a complaint |
| Sales volume rises faster than returns | Denominator inflation | Neutral. Nothing about the experience changed |
Only the top two rows describe a better business. Rows three and four describe a business that has converted a return into resentment. Row five describes a promotion.
All five produce the same arrow.
The counter-intuitive one, and it is well evidenced
The row that surprises most teams is the second, and it comes from the strongest evidence base in this literature: Janakiraman, Syrdal and Freling, "The Effect of Return Policy Leniency on Consumer Purchase and Return Decisions: A Meta-Analytic Review", Journal of Retailing 92(2), 2016. The University of Texas at Dallas summary of the study describes it as covering 21 papers drawn from economics, marketing, decision science, consumer psychology and operations research.
The meta-analysis breaks "leniency" into five separate dimensions, and this is the part almost everyone misses. Leniency is not one lever.
| Leniency dimension | What it controls | Effect found |
|---|---|---|
| Time | Length of the return window | Reduces returns, via the endowment effect |
| Monetary | How much is refunded | Increases purchases |
| Effort | How much hassle the return requires | Increases purchases |
| Scope | Which items may be returned | Part of the overall leniency effect |
| Exchange | Cash refund versus store credit | Part of the overall leniency effect |
Overall, leniency increased purchases more than it increased returns. But the dimension-level result is what matters operationally: lengthening the return window reduces returns, while reducing the effort required to return increases purchases. As one of the authors, Ryan Freling, put it, "In general, firms use return policies to increase purchases but don't want to increase returns, which are costly. But all return policies are not the same."
Set that against the fee trend and the inversion becomes stark. Two opposite interventions both push the return rate down. Extending your window from 30 to 90 days reduces returns while making the purchase safer. Adding a $6 return fee reduces returns while making the purchase riskier. The metric records both as improvement. The first grows the business and the second taxes it, and no amount of staring at the return rate will separate them.
This is the same structural trap we described in choice architecture and defaults: an intervention that improves the headline number by removing the customer's ability to act is reported as a win by the team that owns the number and as a loss, much later, by a team that never sees the connection.
The suppression test
Here is the diagnostic, and it takes one question per intervention.
For any change that lowered your return rate, ask: did it change what the customer knew before buying, or what it cost them after buying?
- Changed what they knew before buying: accuracy. Better sizing guidance, a missing spec added, a photograph that shows scale, a clearer materials list, the return policy actually linked from the product page.
- Changed what it costs them after buying: suppression. Return fees, restocking charges, shorter windows, mandatory in-store drop-off, required photographic proof, refunds paid as store credit.
Accuracy interventions act before the money changes hands and reduce the number of bad matches created. Suppression interventions act afterwards and reduce the number of bad matches you find out about. The bad match still exists in both cases. In the second one it is living in your customer's cupboard, and we follow that population in the regret that never files a return.
The pairing that makes returns reduction falsifiable
A single metric cannot distinguish these. Two can, provided the second one moves in opposite directions under the two mechanisms.
Pair the return rate with the repeat purchase rate of the non-returning cohort, measured at 90 and 180 days.
| Intervention type | Return rate | Repeat purchase rate of keepers | Verdict |
|---|---|---|---|
| Accuracy: added the missing spec | Down | Up or flat | Real improvement |
| Time leniency: window extended | Down | Up or flat | Real improvement |
| Suppression: return fee introduced | Down | Down | Trade, not a win |
| Suppression: window shortened | Down | Down, plus complaint volume up | Trade, and a visible one |
| Denominator: promotion ran | Down | Flat | Artefact. Ignore |
The logic is simple. Accuracy raises the share of purchases that were correct, so the people who kept the item are more satisfied and buy again. Suppression raises the share of customers who kept something they did not want, so the people who kept the item are less satisfied and do not.
Two operational requirements make this work, and both are commonly missing. Tag the cohort with the intervention they were exposed to, or the counterfactual is gone before anyone asks the question. And measure the keepers specifically, not all customers, because blending returners back in reintroduces exactly the population whose behaviour you are trying to hold constant. The cohort mechanics are the same ones described in the cohort analysis guide, and if you are trying to attribute the change causally rather than descriptively, the designs in the quasi-experimental design guide apply directly.
Why nobody catches this
The failure is organisational before it is analytical, and it has the same shape every time.
Different clocks. The return rate responds within one shipping cycle. The repeat purchase effect appears one to two purchase cycles later, which in most categories is two or three quarters. The programme is declared successful and the team is reassigned before the cost lands.
Different owners. Return rate belongs to operations or finance, because returns are a cost line. Repeat purchase belongs to growth or CRM. The suppression cost shows up in a different function's number, in a different quarter, attributed to a different cause, usually "market conditions".
The stated target is the ambiguous metric. In the NRF survey, retailers named reducing return rates as a top priority for 2026, and 64% said they would prioritise updating their returns process within six months. A large share of the industry is about to optimise a number that cannot distinguish serving the customer from deterring them.
The suppressed customer is silent. Someone refused a return complains. Someone deterred by a $6 fee does not, and produces no artefact at all.
How to research this properly
Analytics will not resolve the ambiguity, because both mechanisms produce the same time series. You need to ask people, and you need to ask a specific population.
Interview the keepers, not the returners. The diagnostic population is customers who did not return, split by whether they considered it. The single most valuable question in returns research is "did you think about sending it back, and what stopped you?" A customer who says "I didn't need to, it was right" and a customer who says "it wasn't worth six dollars and a trip to the post office" are both counted as retained.
Interview through a policy change. If you are adding fees or changing windows, that is a natural experiment and the best research moment you will get all year. Sample before and after, ask the same questions, and read the difference.
Ask about the expectation, not the policy. People are poor at predicting their reaction to a hypothetical fee and good at recounting what they actually did. Never ask "would a $6 return fee stop you buying". Ask what happened the last time they wanted to return something and did not.
Quantify inside the conversation. Ranking the deterrents against each other separates a genuine blocker from a mild irritation in a way that rating each one on its own never does.
Where Koji fits
The measurement problem here is that the population you need to hear from is defined by something that never happened. They did not return, did not complain, did not open a ticket. There is no list of them anywhere in your systems except as rows in the orders table marked delivered, which is also where all your happy customers are.
Koji is built for exactly this shape of problem. You define the population, and AI-moderated voice or text interviews run at whatever scale the question needs, on the customer's own schedule, with real-time probing on every vague answer. When someone says the return "seemed like a hassle", the AI asks what specifically, every time, without a researcher in the room.
For this study in particular, Koji's six structured question types do real work alongside the conversation: open_ended, scale, single_choice, multiple_choice, ranking and yes_no, all in one session. A yes_no on whether they considered returning splits the sample. A ranking of the deterrents orders them. A scale captures repurchase intent. A single_choice pins the specific blocker and a multiple_choice captures which policy details they were aware of at purchase. And open_ended captures the account that explains all of it, with automatic thematic analysis clustering the responses and attaching supporting quotes rather than leaving you a spreadsheet of verbatims. See structured questions for how quantitative and qualitative combine in a single study.
The economics are what make a continuous programme possible. Running this as a traditional study means recruiting a screened cohort through a panel, scheduling moderated sessions and paying an analyst to code the transcripts, which is a six-week project with a five-figure cost and no realistic prospect of running it again next quarter. Koji turns it into a study you can rerun every time you change the policy, which is the only cadence at which the before-and-after comparison actually works. No research expertise required to field it, and the report arrives in hours.
What to do in the next quarter
- Classify every returns initiative on the roadmap as accuracy or suppression. Use the before-buying versus after-buying test. Most programmes contain both and report a single number.
- Instrument the pairing before you ship anything. Return rate alone cannot grade the work you are about to do.
- Tag cohorts with the policy they experienced. This is a small engineering task now and an impossible reconstruction later.
- If you are adding fees, capture the baseline first. Both the reason distribution, for the reasons set out in product returns research, and the repeat purchase rate of keepers.
- Run 40 keeper interviews split by considered-returning. The deterred group is the one that does not exist in any other dataset you own.
- Set the reporting rule. A return rate improvement is reported alongside the keeper repeat purchase rate, or it is not reported as an improvement.
Related reading
- Product Returns Research (2026): Why the Reason Code Is a Dropdown, Not a Finding
- Buyer's Remorse Research (2026): The Regret That Never Files a Return
- Choice Architecture and Defaults (2026)
- Product Detail Page Research (2026)
- Cohort Analysis Guide
- Quasi-Experimental Design Guide
Frequently Asked Questions
Is a lower return rate always a good thing?
No. A return rate falls when listings get more accurate, when the return window is extended, when returning becomes costly or difficult, when eligibility is narrowed, and when sales volume outpaces returns. Only the first two describe a better business. The others describe suppression or an accounting artefact, and all five produce an identical improvement on the dashboard.
Does extending the return window increase returns?
The evidence points the other way. The 2016 Journal of Retailing meta-analysis by Janakiraman, Syrdal and Freling found that time leniency, meaning a longer return window, reduces return rates, attributed to the endowment effect as attachment to the product grows over time. Monetary and effort leniency increase purchases. This is why treating leniency as a single dial is a mistake.
Do return fees reduce returns?
They reduce recorded returns. The NRF found merchants charging for at least some returns rose from 66% to 72% in the year the national return rate fell from 16.9% to 15.8%. What a fee changes is whether the customer completes the return, not whether the purchase was wrong. The same survey found 71% of consumers are less likely to shop with a retailer again after a poor returns experience.
How do I tell accuracy from suppression?
Ask whether the change altered what the customer knew before buying or what it cost them after buying. Then pair the return rate with the repeat purchase rate of customers who kept the item, measured at 90 and 180 days. Accuracy pushes that second number up or flat while returns fall; suppression pushes it down while returns fall.
Which customers should I interview about returns reduction?
The keepers, not the returners, split by whether they considered returning. Customers who considered it and were deterred are invisible in every dataset you own, because their defining action is one they did not take. They are counted as satisfied retained customers and they behave like churned ones.
How quickly can we run this study?
With AI-moderated interviews, a keeper study of 40 to 100 customers fields in days rather than the six weeks a panel-recruited moderated study takes. That speed is what makes before-and-after measurement around a policy change practical, since the comparison only works if you can run the same study twice within one policy cycle.
Grade your returns programme before you scale it
If your return rate improved last quarter, there is a version of that story where you served customers better and a version where you taxed them. The dashboard shows both the same way.
Run a keeper study with Koji and find out which one you shipped.