Multi-touch attribution tells you which touchpoints appeared before a conversion. Marketing mix modeling tells you which spend correlates with aggregate sales. Incrementality testing tells you what would have happened if you had not advertised. Only the third is a causal claim, and the gap between it and the other two is not a rounding error: in the best-documented comparison available, an attribution-style estimate put return on ad spend above 4,100 percent while the randomized experiment on the same spend returned minus 63 percent.
That is the whole problem in one line. Below is what each method actually measures, what the peer-reviewed evidence says about how far each can drift from the truth, how the 2026 tooling landscape changed the economics, and the one question that none of the three can answer no matter how well you run them.
The three methods, side by side
| Multi-touch attribution (MTA) | Marketing mix modeling (MMM) | Incrementality testing | |
|---|---|---|---|
| Unit of analysis | The individual user journey | Aggregate spend and sales, usually weekly | A randomized or quasi-randomized group |
| Core technique | Rule-based or algorithmic credit splitting | Regression on time series, increasingly Bayesian | Randomized controlled trial, geo holdout, matched market |
| What it needs | User-level identity across touchpoints | 2-3 years of spend and sales history | Ability to withhold ads from a real group |
| Handles offline and brand media | No | Yes | Yes, with geo designs |
| Survives identity loss | Poorly | Yes | Yes |
| Speed to answer | Near real time | Weeks to months per refresh | Weeks per test, one question at a time |
| Is the output causal | No | Only under strong assumptions | Yes, within the tested population |
| Typical failure mode | Credits demand it did not create | Confounds correlated spend, opaque baseline | Underpowered, narrow, expensive |
What multi-touch attribution actually measures
MTA observes which ads a converting user was exposed to, then divides credit among them by a rule (last click, linear, time decay) or a model. It is a bookkeeping exercise on the population of people who converted. It has no control group, so it cannot distinguish an ad that caused a purchase from an ad that merely found a buyer who was already going to purchase.
That distinction is not theoretical. In March 2012 eBay ran a test that has become the canonical result in this literature: it halted brand-keyword paid search on MSN and measured what happened to traffic. Blake, Nosko and Tadelis reported in Econometrica (83(1):155-174, 2015) that 99.5 percent of the forgone paid click traffic was immediately captured by natural search. Substitution was nearly complete. The ads had been buying clicks that the company would have received for free, and every attribution system in the stack had been recording those clicks as ad-driven conversions.
The same paper ran the harder test on non-brand keywords. Estimated with conventional regression methods of the kind attribution vendors use, return on investment came out above 4,100 percent with no controls and above 1,400 percent with time and geographic controls. The randomized experiment on the same campaigns produced an ROI of minus 63 percent, with a 95 percent confidence interval of [-124 percent, -3 percent]. The sign was wrong, not just the magnitude.
The evidence that "more data" does not fix it
The standard rebuttal is that eBay used crude methods, and that modern attribution with rich user-level features closes the gap. That claim has been tested directly and it does not hold.
Gordon, Zettelmeyer, Bhargava and Chapsky ran the comparison inside Facebook using 15 US advertising experiments comprising 500 million user-experiment observations and 1.6 billion ad impressions (Marketing Science 38(2):193-225, 2019). They estimated each campaign's effect twice: once from the randomized experiment, and once from a battery of matching and regression methods using demographic and behavioral variables far richer than most advertisers ever see.
The results are worth stating precisely, because they are routinely softened in secondary write-ups:
- In half of the studies, the estimated percentage increase in purchase outcomes was off by a factor of three or more across all observational methods.
- In 7 of the 14 studies with a checkout-conversion outcome, the point estimates were consistently off by more than a factor of three.
- Observational methods mostly overestimated the true lift, but not always: in one study all methods except exact matching on age and gender underestimated it, so you cannot even apply a correction factor with a known sign.
- In the most extreme case, the randomized lift was 2.4 percent and the closest observational estimate was 1,306 percent.
Adding variables helped, but never reliably. The authors conclude that commonly used observational approaches based on the data usually available in the industry often fail to accurately measure the true effect of advertising. This is the empirical basis for treating attribution output as a directional operational report, not as evidence.
Why the referee is also underpowered
Incrementality testing is the only method in the set that produces a causal estimate, which makes it tempting to treat as ground truth. It is better than the alternatives, and it is still harder than most teams assume.
Lewis and Rao assembled 25 large digital advertising RCTs accounting for 2.8 million dollars of expenditure, with the median campaign reaching over one million individuals (Quarterly Journal of Economics 130(4):1941-1973, 2015). Their finding is about statistical power, and it is brutal:
- The median standard error on ROI for the retail experiments was 26.1 percent, implying a confidence interval more than 100 percentage points wide. For the brokerage experiments the median standard error was 115 percent.
- The median campaign would have to be nine times larger to reliably distinguish a wildly profitable campaign (+50 percent ROI) from one that merely broke even.
- To resolve a 10 percent ROI difference, the median campaign would have to be 62 times larger.
The mechanism is variance, not bad design. Individual-level sales are enormously volatile relative to the effect an ad needs to have: the standard deviation of individual sales is typically about ten times the mean. To net a 25 percent ROI, the median campaign in their sample had to raise average per-person sales by 35 cents, on a variable with a mean of 7 dollars and a standard deviation of 75 dollars. The implied model fit for a highly profitable campaign is an R-squared on the order of 0.0000054.
Read that number again. A campaign good enough to be one of the best investments the company made that quarter explains roughly five millionths of the variance in sales. That is the signal every measurement method in this article is trying to recover.
What changed in 2025 and 2026
Three things reshaped the practical landscape, and they pushed in the same direction.
Open-source MMM collapsed the cost of entry. Google released Meridian to everyone on 29 January 2025, after testing with hundreds of brands. It uses Bayesian causal inference and, critically, it integrates incrementality experiment results as priors, agnostic of the channel or the experiment - the model is designed to be calibrated by experiments rather than to replace them. Meta's Robyn had already established the open-source pattern from the other direction, using ridge regression with evolutionary hyperparameter optimization. MMM stopped being a six-figure consulting engagement and became a library.
The identity crisis resolved in an unexpected direction. On 22 April 2025 Google announced it would "maintain our current approach to offering users third-party cookie choice in Chrome, and will not be rolling out a new standalone prompt for third-party cookies." Chrome kept third-party cookies. Later in 2025 Google went further and began retiring most of the Privacy Sandbox APIs that were supposed to replace them. The reprieve is real but partial: it does nothing about Apple's App Tracking Transparency, walled-garden reporting, or consent-mode gaps in the EU. Teams that rebuilt around aggregate and experimental measurement were not wrong; teams that waited for cookies to die and did nothing lost several years.
Adoption followed. Marketers have moved toward running MMM and MTA in parallel rather than choosing, with incrementality tests used to calibrate both. That layered design is the right instinct. It also multiplies the number of ways a stack can produce a confident number with no causal content.
What each method can and cannot prove
| Question you are asking | MTA | MMM | Incrementality | Customer research |
|---|---|---|---|---|
| Which creative should I pause today | Usable | No | Too slow | No |
| How should I split next year's budget | No | Yes | Partially | No |
| Did this channel cause sales | No | Weakly | Yes | No |
| Does brand and offline media work | No | Yes | Yes, with geo | No |
| Why did the campaign work or fail | No | No | No | Yes |
| What almost stopped the purchase | No | No | No | Yes |
| What would have persuaded the people who did not convert | No | No | No | Yes |
The right-hand column is the point of this article. Every method in the first three columns is a form of arithmetic on outcome data. They differ in how honestly they handle the counterfactual, and incrementality handles it best. None of them contains any information about mechanism, because none of them ever asks a person a question.
This shows up most obviously in the baseline. In a typical MMM, the largest single term is base sales - the volume the model attributes to everything that is not measured media. It is frequently the majority of the modelled outcome, and it is the part of the business the model has the least to say about. You can decompose the media tail with Bayesian rigor and still have no account of the largest number on the chart. Much of that baseline is demand created by category entry points - the situations that make someone think of the category at all, which originate outside your media plan entirely.
The question none of the three can answer
An experiment tells you that removing a channel cost you 4 percent of conversions. It does not tell you:
- Which of the several things the ad communicated was the one that mattered
- What the buyer was actually trying to accomplish when the ad reached them
- What nearly stopped them, and what resolved it
- What the people who saw the ad and did not buy were thinking
- Whether the effect will survive a message change, a price change, or a new competitor
Those are questions about reasons, and reasons live in people, not in aggregates. Historically the reason teams accepted this gap is that closing it meant commissioning qualitative research: recruiting, scheduling, moderating, transcribing and coding dozens of conversations over six to eight weeks, by which time the budget decision was made. The measurement number arrived in a day and the explanation arrived in two months, so only one of them was ever in the room.
Equalize the latency and the two get weighed together. That is what changed.
How Koji closes the mechanism gap
Koji runs AI-moderated voice interviews at survey scale. You write a brief, Koji generates the discussion guide, and the AI consultant conducts every conversation - probing follow-ups included - with hundreds of respondents in parallel. Thematic analysis is automatic, and the report is one click.
For measurement teams specifically:
- Ask the people your experiment cannot explain. Run a holdout, then interview both cells about the same purchase. The experiment gives you the size of the effect; the interviews give you its mechanism.
- No moderator bias across the sample. A human moderator running interview 40 asks different questions than they did in interview 1. The AI consultant asks the same core questions of everyone, then probes on what each person actually said - so your qualitative data supports quantitative comparison instead of undermining it.
- Quantitative and qualitative in one instrument. Koji supports six structured question types alongside open conversation:
open_ended,scale,single_choice,multiple_choice,ranking, andyes_no. You can size a segment and hear it explain itself in the same study. See the structured questions guide for how each type is analyzed and charted. - Hours, not weeks. From question to insight in the same cycle as the media decision, which is the only reason the explanation ever makes it into the meeting.
Useful companions in the docs: A/B testing vs user research on what experiments can and cannot tell you, quasi-experimental design for when randomization is impossible, Koji for marketing teams for campaign workflows, ad testing surveys for pre-flight creative work, and customer journey interviews for mapping what happened around the purchase.
What to actually do
- Stop treating attributed ROAS as evidence. It is an operational dashboard. Use it to catch outages and pacing problems, not to defend a budget.
- Run incrementality tests on your largest and most suspicious line items first. Branded search and retargeting are where substitution is highest and the eBay result applies most directly. That is doubly true inside a retail media network, where the party selling the ads is also the party reporting the result.
- Size your test before you run it. If your campaign cannot detect a 50 percent ROI difference, do not report a point estimate as though it can.
- Calibrate MMM with experiments rather than arguing about which is right. Meridian is built for exactly this, and it is free.
- Attach a mechanism study to every material test. The experiment tells you whether to keep spending. The interviews tell you what to change.
The teams that get this right in 2026 are not the ones with the most sophisticated model. They are the ones who stopped asking their measurement stack a question it was never built to answer, and started asking their customers instead.
Ready to find out why your numbers moved? Start a study with Koji and run your first AI-moderated interviews today - no research background required, and results in hours instead of weeks.
Frequently Asked Questions
Is marketing mix modeling better than multi-touch attribution?
For budget allocation, yes. MMM sees every channel including offline and brand media, does not depend on user-level identity, and is not distorted by the walled-garden reporting that MTA inherits. But "better" is doing a lot of work: MMM is still a regression on correlated spend, so it is only causal under strong assumptions. The rigorous version of the answer is that MMM is the better strategic instrument and incrementality testing is the only causal one. Use experiments to calibrate the model, which is what Google Meridian is explicitly designed to support.
How wrong can attribution actually be?
Measurably and unpredictably wrong. Gordon and colleagues found that in half of their 15 Facebook experiments, observational estimates of the increase in purchases were off by a factor of three or more, with one case comparing a true 2.4 percent lift to a 1,306 percent estimate. Blake, Nosko and Tadelis found eBay's non-brand search ROI estimated at over 4,100 percent by regression and minus 63 percent by experiment. The errors do not have a consistent direction, so you cannot correct for them with a fudge factor.
What is incrementality testing and how is it different?
Incrementality testing withholds advertising from a randomly chosen group and compares outcomes against an exposed group. Because assignment is random, the difference between the groups is caused by the advertising rather than merely associated with it. Common designs include user-level RCTs, geo holdouts and matched-market tests. It is the only one of the three methods that estimates the counterfactual directly - what would have happened without the spend.
Do I still need this now that Chrome kept third-party cookies?
Yes. Google's 22 April 2025 decision removed one source of signal loss, not the category. App Tracking Transparency, walled-garden reporting, cross-device gaps and EU consent requirements all persist, and none of them are what makes attribution non-causal in the first place. The eBay and Facebook results predate cookie deprecation entirely. Perfect tracking would not have fixed either one, because the problem is the missing control group, not the missing identifier.
How large does an incrementality test need to be?
Larger than most teams expect. In Lewis and Rao's sample of 25 RCTs, the median campaign - already reaching over a million people - would have needed to be nine times larger to reliably separate a +50 percent ROI campaign from a break-even one, and 62 times larger to resolve a 10 percent difference. The reason is that individual sales are roughly ten times more variable than their mean, while the effect an ad needs to produce is small. Always run a power calculation before you commit, and be willing to report "we could not tell" as a result.
Can customer research replace incrementality testing?
No, and it should not try to. They answer different questions and fail in different places. An experiment measures the size of an effect but is silent on why it happened; interviews recover reasons, mechanisms and objections but cannot size an effect in the population. The productive pattern is to pair them: run the holdout to learn whether the spend works, then interview both cells to learn what it did to the people in them. With Koji that second step takes hours rather than the six to eight weeks that made teams skip it.