Back to blog
Research

Product Returns Research (2026): Why the Reason Code Is a Dropdown, Not a Finding

Return reason codes measure your refund policy, not the cause of the return. Here is why the data is biased, what it cannot distinguish, and how to interview returners about the purchase instead.

Koji

Koji Team

Research · · 13 min read

A return reason code is not a research finding. It is a dropdown, written by you, answered by a customer whose only goal at that moment is to get their money back quickly and cheaply. The distribution you get back is sorted by what each option costs the person selecting it, not by what actually went wrong. That is why "wrong size" and "arrived damaged" dominate almost every returns dataset in retail, and why the categories that would actually change your roadmap barely register. Fixing this does not require better codes. It requires talking to returners, which platforms like Koji make practical at a scale that used to need a research team.

Returns are now the largest, best-funded and least-researched dataset most retailers own. The money is enormous, the operational attention is intense, and the causal understanding is close to zero.

The short answer

Return reason codes tell you what a customer was willing to declare in order to get a refund approved. They do not tell you why the purchase failed. The purchase failed earlier, on a page you wrote, against an expectation you never measured. To find the cause you need to interview returners about the moment of purchase, not the moment of return, and you need to do it while the memory is intact.

What the 2026 numbers actually say

The best primary source in this space is the 2025 Retail Returns Landscape, published 15 October 2025 by the National Retail Federation with Happy Returns, a UPS company. It surveyed 2,006 consumers who returned at least one online purchase in the previous 12 months, plus 358 ecommerce professionals at US merchants with more than $500 million in revenue, in summer 2025.

The headline figures:

  • $849.9 billion in merchandise expected to be returned in 2025, equal to 15.8% of annual sales, down from 16.9% and $890 billion in 2024.
  • 19.3% of online sales returned, against 17% of holiday sales.
  • 9% of all returns are fraudulent. Among retailers that track it, the reported increases are overstated quantity of returns (71%), empty box or "box of rocks" returns (65%) and decoy returns such as counterfeit items (64%).
  • 82% of consumers say free returns are an important consideration, up from 76% the year before, and 76% prefer an instant refund or exchange.

One number in that report deserves far more attention than it gets. 45% of shoppers say it is acceptable to "bend the rules" on a return. That is not a fraud statistic. Fraud is the 9%. This is something different and more damaging to your data: nearly half your returning customers consider misstating a return to be socially acceptable behaviour, and the place they misstate it is the reason field.

Treat that as what it is. It is a direct, primary-sourced measurement of the unreliability of the exact instrument the entire returns industry uses to diagnose itself.

A caution on sourcing, because this topic attracts it: search for return rate benchmarks by category and the first page is almost entirely content farms publishing confident category-level percentages with no primary source anywhere. Several contradict each other by a factor of two. Some restate the NRF online figure as 24.5% when the report says 19.3%. Use the NRF report, use Baymard, and ignore the rest.

The instrument problem, stated plainly

Every return reason dataset shares four properties, and each one is a measurement defect:

You wrote the options. The customer can only report a cause you already thought of. A reason that is not on the list is invisible by construction, which means the dataset can never surface a cause you have not already guessed. This is the returns version of a closed-ended survey with no "other" box, and it fails the same way.

They answer at the wrong moment. The code is collected at the return, which is days or weeks after the purchase. The decision that went wrong happened at the listing. By the time someone opens the returns portal they have reconstructed the story, and what they type is an account of the outcome, not the decision.

They answer under a cost gradient. This is the important one. The options are not equally priced.

Nobody validates it. The reason code is never checked against the item, the listing or the customer. It goes straight into a dashboard as though it were an observation.

The refund-cost gradient

Here is the mechanism that shapes almost every returns dataset in retail. Return reasons carry different consequences for the person selecting them, and customers learn the gradient fast.

Reason selectedTypical consequence for the customerDirection of bias
Arrived damaged or defectiveFree return, fast refund, no fee, merchant absorbs shippingHeavily over-selected
Wrong item sentFree return, priority handling, no feeOver-selected
Doesn't fit / not as describedUsually free, occasionally a feeOver-selected
Changed my mindReturn fee or restocking charge likely, slower refundHeavily under-selected
No longer needed / bought elsewhereFee likely, sometimes refused outside windowHeavily under-selected
Bought several, keeping oneRarely offered as an option at allEffectively unmeasurable

Read the right-hand column as a single sentence: your returns data is biased toward blaming you and away from blaming the customer, in direct proportion to how much you charge the customer for admitting fault.

That yields a testable prediction, and it is the most useful thing in this article. The moment you introduce or raise return fees, your returns data becomes less accurate, not more. The fee does not change why people return things. It changes which box they tick. Fault-of-merchant codes rise, fault-of-customer codes fall, and a returns team reading the dashboard concludes that product quality has deteriorated when nothing about the product has changed at all. We take that consequence apart in why a falling return rate is two different stories.

What a reason code cannot distinguish

The deeper problem is that a single code collapses causes with completely different owners and completely different fixes.

The code saysIt could meanWho should act
Doesn't fitSize chart is wrongMerchandising
Doesn't fitSize chart is right, the photography implied a different cutCreative
Doesn't fitCustomer ordered two sizes deliberatelyNobody, this is planned
Not as describedThe copy overclaimedCopy and legal
Not as describedThe copy was accurate but the main image was notCreative
Not as describedThe customer read a competitor's spec and assumed it appliedCategory strategy
Changed my mindGenuine impulse reversalNobody
Changed my mindFound it cheaper elsewhere during the delivery windowPricing
Changed my mindDelivery took so long the occasion passedLogistics

Nine distinct causes, four teams, three codes. No amount of dashboard filtering recovers the distinction, because the distinguishing information was never collected. This is the same structural failure we describe in cart abandonment research: the analytics tell you where, and only a conversation tells you why.

The pre-purchase gap that produces the return

Returns research fails when it stays inside the returns funnel. The cause is upstream, and there is good evidence about where.

Baymard Institute's product page benchmark, drawn from 30,000+ manually scored product page usability assessments across 155+ benchmarked sites, finds that only 48% of desktop, 38% of mobile and 36% of app product pages score "decent" or "good". The specific failures are almost all information failures, and two of them bear directly on returns: 44% of product pages do not display or link to the return policy from the main product page content, and 67% give no estimated total order cost near the buy section.

That is the return being manufactured. A shopper who cannot find the return policy either does not buy, or buys with an assumption about what returning will cost them. When the assumption is wrong, you get a return and a poor experience at once. The listing generated both, and the reason code will record neither.

If you want the full argument about what product page tests can and cannot see, we wrote it up in product detail page research.

How to research returns properly

The method is not complicated. It is just different from what returns teams currently do.

1. Sample from returners, not from return reasons. Do not stratify by code. The codes are the contaminated variable. Stratify by product category, order value and time since delivery.

2. Interview about the purchase, not the return. The first question is never "why did you return it". It is "walk me through deciding to buy this". You are trying to recover the expectation, because a return is always a gap between an expectation and an object.

3. Ask what they expected the item to be. Then ask what it turned out to be. The difference is the finding. This one pair of questions recovers more actionable detail than an entire year of reason codes.

4. Ask where the expectation came from. The listing, a photograph, a review, a competitor's page, a friend, or an assumption carried in from another category. Only the first three are things you control, and knowing which one it was tells you exactly which team owns the fix.

5. Ask what would have stopped the purchase. Returners are the only population who can answer this concretely, because they know how the story ended.

6. Time it to the return, not to the refund. Memory of the purchase decision decays fast. Reach people within days of the return being initiated, while both ends of the gap are still available to them.

Returners are the most under-used sample in ecommerce for a simple reason: they are pre-qualified, recently active, contactable, and they have already demonstrated that something went wrong by paying a cost to tell you. Roughly one in five online orders coming back is a listing accuracy statistic before it is a logistics statistic.

Where Koji fits

The reason returns research is rare is not that anyone thinks it is a bad idea. It is that the population is large, geographically scattered, mildly annoyed with you, and unlikely to accept a scheduled video call from the company they just returned something to. Traditional research economics do not work here. Recruiting 200 returners through a panel, scheduling moderated sessions and having a researcher analyse the transcripts is a six-week, five-figure project, and by the time it lands the assortment has moved on.

Koji removes the constraint that makes this hard. You write the brief, Koji runs AI-moderated voice or text interviews with returners on their own schedule, and the follow-up questions happen in the conversation rather than in a researcher's calendar. When a customer says "it wasn't what I expected", the AI asks what they expected and where that expectation came from, in the moment, every single time. That is the probe a dropdown cannot perform and a busy human moderator performs inconsistently.

Three things make it fit this problem specifically:

  • Structured questions alongside open conversation. Koji supports six question types in a single study: open_ended, scale, single_choice, multiple_choice, ranking and yes_no. That matters here because you want both the quantified read (rank these five factors in how much they contributed to sending it back) and the unprompted narrative (what did you think you were buying), and you want them from the same person in the same session. See structured questions for how the two combine in one study.
  • Automatic thematic analysis. Returns research generates the kind of messy, overlapping, category-specific language that makes manual coding expensive. Koji clusters it into themes with the supporting quotes attached, so the finding arrives with its evidence rather than as a percentage you have to trust.
  • No moderator bias. Every returner gets the same opening and the same probing logic. Nobody gets a leading question because the moderator already has a theory by interview forty.

The practical difference is timeline. This is a question-to-insight-in-hours workflow rather than a six-week study, which means returns research can run continuously against a moving assortment instead of once a year against a snapshot.

If you also want the transactional side instrumented, post-purchase surveys cover the satisfaction read, and Koji handles the diagnostic conversation that a survey cannot.

What to do in the next quarter

  • Audit your reason list against the cost gradient. For each option, write down what selecting it costs the customer. If the cheap options dominate your distribution, you have measured your policy.
  • Add a genuine free-text field and actually read it. Not as the primary instrument, but as a source of the categories missing from your list.
  • Run 40 returner interviews in one category. One category, not all of them. The findings are category-specific and mixing them produces mush.
  • Diff the expectation against the listing. For each finding, open the product page and locate the sentence that should have prevented it. Sometimes it is missing. Sometimes it is present and unread, which is a different fix.
  • Check whether your product pages link the return policy. If Baymard's 44% is representative, there is a coin-flip chance yours does not.
  • Re-baseline before and after any fee change. If you are introducing return fees this year, capture your reason distribution first. Otherwise you will read the tick-box migration as a product quality change.

Frequently Asked Questions

Why are return reason codes unreliable?

Because the customer selects from a list you wrote, at a moment when their goal is a fast, cheap refund. Options that are free and blame the merchant get over-selected; options that trigger a fee or blame the customer get under-selected. The NRF's 2025 report found 45% of shoppers consider it acceptable to bend the rules on a return, which is a direct measurement of that unreliability. The resulting distribution describes your refund policy more faithfully than it describes any cause.

What is the average ecommerce return rate in 2026?

The most defensible primary figure is the NRF and Happy Returns 2025 Retail Returns Landscape: 15.8% of total US retail sales and 19.3% of online sales, amounting to $849.9 billion, down from 16.9% and $890 billion in 2024. Be sceptical of category-level benchmarks circulating online, as most trace to no primary source and frequently contradict each other.

Should I interview customers about the return or about the purchase?

About the purchase. A return is the gap between what someone expected and what arrived, so the diagnostic information sits at the moment the expectation formed, which is the listing. Asking why they returned it yields a restatement of the reason code. Asking what they thought they were buying, and where that impression came from, yields the fix and its owner.

How soon after a return should I run the research?

Within days of the return being initiated. Memory of the purchase decision decays quickly, and the specific detail you need, which photograph they looked at or which line of copy they read, is the first thing to go. Waiting for the refund to settle typically costs you the most valuable part of the account.

Do return fees improve returns data quality?

No, they degrade it. Fees change which reason a customer selects without changing why the item came back, pushing declarations toward the free, no-fault options. Teams then read the shift as a product quality problem. If you are changing fees, capture a reason-code baseline first so you can separate the tick-box migration from any real change.

Can AI interviews handle a population as large as returners?

Yes, and this is the case where AI moderation matters most. Returners are numerous, scattered and unwilling to schedule a call with a company they just returned something to. Koji runs AI-moderated voice or text interviews on the customer's own schedule, probes every vague answer in the moment, and produces themed analysis with supporting quotes, turning a six-week panel study into a continuous programme.

See what your returns data cannot tell you

Your return reasons are a record of your refund policy. The cause of the return is still sitting in the memory of the person who sent the item back, and it is recoverable for about a week.

Start a returns study with Koji and hear what your customers thought they were buying.

Run your first AI-moderated study in 10 minutes

10 free credits on signup. No credit card required.

GDPR compliantEU or US data residencyNo AI training on your data
Koji

Koji Team

Research

Share this article

Keep reading