In retail media, the company selling you the advertising is also the company that defines what counts as an exposure, sets the attribution window, owns the only record of the purchase, and publishes your return on ad spend. No other media channel concentrates all four of those roles in one party. That is not an accusation of bad faith. It is a structural description, and it means a network-reported ROAS is a vendor's self-assessment rather than an independent measurement.
The stakes are now large enough that the distinction is expensive. eMarketer's 2026 forecast puts US retail media ad spend at 71.09 billion dollars, roughly 18 percent growth year over year, with about 89 percent of the 10.53 billion dollars in incremental investment going to Amazon and Walmart alone. Amazon's own advertising revenue passed 68 billion dollars across 2025, with 21.3 billion dollars in the fourth quarter, up 22 percent year over year - a business still largely fueled by sponsored product listings on its marketplace.
This article covers what the industry standards actually require, the four specific reasons reported numbers overstate causal effect, and what to do instead.
What the IAB/MRC guidelines actually say
The IAB and the Media Rating Council released the Retail Media Measurement Guidelines for public comment on 13 September 2023 and finalized them at IAB's Annual Leadership Meeting on 29 January 2024. They are the closest thing the category has to a standard, and reading them carefully is more useful than reading any vendor's methodology page.
Three requirements matter most.
Attribution windows must be justified, not chosen. The guidelines state that lookback windows "must be empirically supported" and "must be logical with regard to campaign objectives and length as well as category sales cycles," and that measurement providers must establish "empirically supported limits to the length of a lookback period... with defensible audit evidence maintained supporting them." Windows must also "be disclosed up front before campaign execution and measurement."
But windows are not standardized across networks. The guidelines concede this directly: "While attribution windows may differ across retail media organizations, consistent granularity of component days is encouraged to allow reconciliation across differing windows." Encouraged, not required.
Reported sales frequently include modeled sales. On closed-loop attribution, the guidelines note that it "often involves models that involve the use of extrapolation techniques to show the full impact of a campaign on brand sales," because not every ad served can be traced to an identified user. Extrapolation is legitimate and disclosed. It is also not the same thing as counting.
Four reasons the reported number overstates the effect
1. The window is a lever, and it is the network's lever
A 14-day window and a 30-day window produce materially different ROAS on identical media, because a longer window sweeps in more purchases that would have happened anyway. Since each network sets its own, cross-network ROAS comparison is invalid by construction - you are comparing two numbers computed under different definitions, from two parties with an interest in the result. When a network reports a stronger ROAS than its competitor, the first question is not "why does this channel work better" but "what window is each one using."
2. Sponsored products intercept demand at the point of maximum intent
Retail media's core inventory sits on search results and product pages - which is to say, in front of people who have already opened a shopping app and typed a query. That is the highest-intent moment in the entire funnel, and it is precisely the situation where advertising is most likely to be credited for a purchase it did not cause.
The canonical evidence remains eBay's. When it halted brand-keyword paid search on MSN in March 2012, Blake, Nosko and Tadelis found that 99.5 percent of the forgone paid click traffic was immediately captured by natural search (Econometrica 83(1):155-174, 2015). Substitution was nearly complete: the ads were buying traffic the company would have received for free.
The same paper measured non-brand search two ways. Conventional regression methods produced ROI above 4,100 percent with no controls and above 1,400 percent with time and geographic controls. The randomized experiment returned minus 63 percent, with a 95 percent confidence interval of [-124 percent, -3 percent]. Retail media search inventory is structurally closer to the brand-keyword case than the non-brand case, which is the uncomfortable part.
3. The incrementality chapter excludes the halo
The IAB/MRC incrementality guidance is genuinely good on method - it treats experimental design as the foundation, and calls for disclosure of control group construction, holdout periods and limitations. But note its stated scope: "The primary focus of this incrementality chapter is on measuring the direct incremental impact within a retailer's ecosystem, rather than considering the broader halo effect it may have on the rest of the business."
That cuts both ways, and brands usually only notice one of them. Sales your retail media drove at other retailers are excluded, which understates the campaign. Sales you would have made at that same retailer anyway are what the incrementality test is supposed to strip out. The net direction depends entirely on which effect dominates in your category - which is another way of saying you cannot infer it from the reported number.
4. Observational measurement does not become causal at scale
The standard defence of closed-loop retail media measurement is that the data is deterministic: real logged-in shoppers, real transactions, no modelling. That defence conflates data quality with causal identification. Gordon, Zettelmeyer, Bhargava and Chapsky tested exactly this claim using 15 advertising experiments comprising 500 million user-experiment observations and 1.6 billion ad impressions (Marketing Science 38(2):193-225, 2019). Against randomized ground truth, observational methods were off by a factor of three or more in half the studies, and in one case compared a true 2.4 percent lift against a 1,306 percent estimate. Richer variables helped inconsistently, and the errors ran in both directions.
Clean data with no control group is still not a causal estimate.
What to ask every retail media network
| Question | What a good answer looks like | Red flag |
|---|---|---|
| What is the attribution window, by campaign type | A specific number, disclosed before launch, with the empirical basis | "It varies" or a number you learn after the fact |
| Is it view-through, click-through, or both | Reported separately, with view-through broken out | A single blended ROAS |
| What share of reported sales is extrapolated | A percentage, with the extrapolation method described | The question is treated as unusual |
| Can we run a holdout | Yes, with a documented design | Only network-run lift studies |
| Who constructs the control group | An independent party, or a design you can audit | The network, unexplained |
| Are new-to-brand buyers reported | Yes, defined and separated from repeat | Only total attributed sales |
| Can we get log-level data | Yes, via a clean room or export | Aggregated dashboards only |
If a network cannot answer the first three, its ROAS is not a number you can put in a budget defence.
What to do instead
Run your own incrementality test. A geo holdout or matched-market design that you or an independent partner controls is the only way to get a causal estimate of a retail media line item. Test the largest and most suspicious items first - branded search and retargeting, where substitution is highest. And size the test before you run it, because advertising experiments are chronically underpowered relative to the sales volatility they fight against.
Compare networks on incrementality, never on ROAS. ROAS is defined differently by each network. Incremental sales per dollar, measured by your own experiment under one definition, is comparable.
Separate demand capture from demand creation in the budget. Sponsored products at the bottom of the funnel is a distribution cost that mostly harvests existing intent. It is not the same activity as building the memory links that make people search for you in the first place - see category entry points for how that demand actually originates. Blending both into a single ROAS number is how brands end up optimizing themselves into pure harvesting, and then wondering why the harvest shrinks.
Then ask the shopper. This is the part no measurement design covers. An experiment can tell you a campaign produced 3 percent incremental units. It cannot tell you whether the ad changed a mind, confirmed a decision already made in the aisle, or simply put your label on a purchase that a rival would otherwise have taken. Those three cases require completely different responses, and they are indistinguishable in transaction data. For the broader picture of what each measurement method can prove, see MMM vs attribution vs incrementality.
How Koji answers the question the logs cannot
Koji runs AI-moderated voice interviews at survey scale, which makes shopper-side evidence fast enough to sit alongside a media decision instead of arriving a quarter late.
- Interview both cells of your holdout. The experiment gives you the size of the effect. Interviewing exposed and unexposed shoppers about the same purchase gives you its mechanism - what they were trying to buy, what they considered, what the ad did or did not change.
- Reach the shoppers who did not convert. They are absent from every retail media report by construction, and they are where the diagnostic information lives.
- Quantify and explain in one instrument. Six structured question types run alongside open conversation:
open_endedfor the purchase narrative,scalefor consideration and intent,single_choiceandmultiple_choicefor brand and retailer sets,rankingfor what actually drove the choice, andyes_nofor clean qualification. The structured questions guide explains how each is analyzed. - No moderator bias, no six-week synthesis. Every respondent gets the same core questions with adaptive follow-ups, thematic analysis is automatic, and the report is one click.
Further reading in the docs: AI research for retail, AI research for CPG, AI research for ecommerce, cart abandonment research, A/B testing vs user research, and Koji for marketing teams.
The short version
Retail media is a genuinely valuable channel that happens to come with a measurement system built by the seller. The IAB and MRC did real work in 2024 to make that system auditable, and the guidelines are worth holding networks to. But standards for disclosure are not the same as independence, and no disclosure regime turns an observational number into a causal one.
Treat network-reported ROAS as inventory telemetry. Get your causal estimate from an experiment you control. Get your explanation from the shoppers themselves. The brands that do all three will spend the next few years reallocating budget away from the line items that were only ever taking credit.
Want to know what your retail media actually changed? Start a study with Koji and interview real shoppers about real purchases - results in hours, no research background required.
Frequently Asked Questions
Is retail media ROAS reliable?
As an operational signal, yes - it tells you whether campaigns are delivering and pacing. As evidence of causal effect, no. The network defines the exposure, sets the attribution window, owns the transaction record and publishes the result, and reported sales often include extrapolated figures because not every impression can be tied to an identified shopper. Use ROAS to run campaigns; use an experiment you control to decide whether the spend is worth making.
Why can I not compare ROAS across retail media networks?
Because each network sets its own attribution window, and the IAB/MRC guidelines only encourage rather than require consistency across organizations. A 30-day window will report a higher ROAS than a 14-day window on identical media, since it captures more purchases that would have occurred anyway. Two ROAS figures computed under different definitions are not comparable. Compare incremental sales per dollar from a single experimental design instead.
What do the IAB/MRC retail media guidelines require?
The guidelines, finalized on 29 January 2024 after a public comment draft in September 2023, require that attribution and lookback windows be empirically supported, defensible under audit, and disclosed before campaign execution. They require transparency about incrementality methodology including control group construction and holdout periods, and disclosure of extrapolation used in closed-loop attribution. They explicitly scope the incrementality chapter to impact within the retailer ecosystem, excluding halo effects elsewhere.
Are sponsored product ads incremental?
Often much less than reported, because they appear in front of shoppers who have already searched for the product. The closest rigorous evidence is eBay's brand-keyword experiment, where 99.5 percent of paid click traffic was recovered by natural search once the ads were switched off. In the same research, non-brand search ROI was estimated above 4,100 percent by regression but measured at minus 63 percent by experiment. The only way to know for your own campaigns is a holdout test.
How do I run a retail media incrementality test?
Withhold the media from a randomly assigned group - typically a set of geographies, since user-level holdouts inside a retailer ecosystem are rarely available to the advertiser - and compare sales against a matched exposed group. Define the metric and window before launch, run a power calculation so you know what size of effect you can actually detect, and use an independent party to construct the control group where possible. Start with branded search and retargeting, where substitution is highest.
What can customer research tell me that an incrementality test cannot?
Mechanism. An experiment can establish that a campaign produced incremental units, but not whether the ad changed a decision, confirmed one already made, or simply captured a purchase a competitor would have taken. It also cannot reach the shoppers who saw the ad and did not buy, who are missing from every retail media report by construction. Interviews recover the reasoning behind the behaviour, and with AI-moderated interviews that evidence arrives in hours rather than weeks.