Incrementality Testing vs MTA vs MMM for Ecommerce
Three measurement methods, three different questions, and one dependency all of them share. A practical guide to choosing between them — and to the event-data problem that silently biases all three.

Every few years the measurement conversation resets. Multi-touch attribution was going to give us per-customer truth. Then identifier loss made that harder, and marketing mix modeling came back from the 2000s wearing new clothes. Now incrementality testing is ascendant, and the vendor pitching each one will tell you the other two are obsolete.
They are not competing answers to one question. They are answers to three different questions, operating at different levels of resolution, with different failure modes. Most ecommerce teams above a certain spend need more than one.
What almost nobody selling these methods will tell you is that all three inherit their accuracy from the same upstream source — the conversion event stream — and that a broken event pipeline biases all three in ways their outputs do not reveal. That part comes later, because it only makes sense once the three methods are clear.
The three methods, precisely
The terms get used loosely enough that it is worth being exact.
Multi-touch attribution
MTA is identity-based. It attempts to observe the sequence of touchpoints an individual encountered before converting, then assigns fractional credit across them according to some rule — linear, time-decay, position-based, or algorithmic.
Its defining property is resolution. MTA is the only one of the three that can tell you anything about a specific customer journey. Its defining requirement is that you can actually link touchpoints to individuals across sessions, devices, and platforms.
Marketing mix modeling
MMM is aggregate and statistical. It takes time-series data — spend by channel, conversions, plus external factors like seasonality, promotions, and pricing — and fits a model estimating each channel's contribution. It never observes an individual.
That is precisely why it survived the identifier collapse untouched. MMM never depended on tracking anyone, so nothing about ATT, ITP, or cookie policy degraded it. Its costs are different: it needs substantial historical data, it produces channel-level rather than campaign-level guidance, and it answers slowly.
Incrementality testing
Incrementality testing is causal and experimental. You deliberately withhold advertising from a randomly selected group — a set of geographies, or an audience segment — run the campaign to everyone else, and measure the difference in outcomes between the two.
This is the only method of the three that establishes causation rather than correlation. It answers the question the other two can only approximate: what would have happened anyway? Its costs are real — you are knowingly forgoing revenue in the holdout, tests take weeks, and each test answers one narrow question.
MTA is weakened, not dead
The obituaries are overstated, and the reason matters for how you use it.
MTA's problem is that identity fragmentation broke its core requirement. Safari's tracking prevention truncates the cookies that link sessions. ATT limits cross-app observation. Click identifiers are increasingly stripped in some contexts. The result is not that MTA produces wrong answers uniformly — it is that MTA now observes a biased subset of journeys and reports on it as though it were the whole.
That bias has a direction. The touchpoints that remain observable skew toward the platforms and contexts where tracking still works well. Channels that are easy to measure get credited for conversions that channels harder to measure actually influenced. This is the over-crediting problem: not that MTA invents conversions, but that it systematically reassigns credit toward whatever it can still see.
You will find confident percentages quantifying that over-crediting in secondary marketing coverage. We could not locate the original research behind the figure most often cited, so we are not going to repeat a number we cannot source. The mechanism is well established; the magnitude is specific to your channel mix and your tracking quality, and honestly the only way to know yours is to test it — which is what incrementality is for.
What the adoption data shows
Incrementality testing has moved from specialist technique to mainstream practice. EMARKETER data from July 2025, published in November 2025, found that 52.0% of US brand and agency marketers use incrementality testing or experiments to measure campaigns. In the same research, 27.6% named expanding incrementality testing a top measurement priority, and 36.2% planned further investment over the following twelve months.
Some context on why the shift happened when it did. Google officially retired the remaining Privacy Sandbox APIs on October 20 2025, including the Attribution Reporting API — the mechanism that had been positioned for years as the privacy-durable replacement for third-party-cookie attribution. Its cancellation removed the option many teams had been waiting on, which left experimentation as the remaining path to causal confidence.
A framework for choosing
The three methods operate on different decision cycles, and matching each to the decisions it can actually inform is most of the work.
MMM for budget allocation across channels, quarterly or annually. When the question is "how much should go to paid social versus search versus retail media next quarter," MMM is the right instrument. It is privacy-durable, it captures offline and brand effects the other two miss entirely, and quarterly cadence suits its slow answer. Do not ask it which creative to pause.
Incrementality testing to validate scaling decisions. Before you double spend on a channel, or when a channel's reported ROAS looks too good, run a holdout. This is the instrument for high-stakes, low-frequency decisions where being wrong is expensive. It is too slow and too costly to run continuously across everything.
MTA for daily tactical optimization inside already-validated channels. Once incrementality has confirmed a channel genuinely drives incremental revenue, MTA is useful for the within-channel questions — which campaigns, which audiences, which creative. Use it for relative comparisons between similar things measured the same way, where its biases apply roughly equally and partly cancel. Do not use it to compare across channels, which is exactly where its bias is worst.
The through-line: the more consequential and less reversible the decision, the more you should weight causal evidence over observational.
Designing a holdout that means something
A badly designed incrementality test produces a confident number that is noise. A few fundamentals worth respecting.
Geo holdouts split markets into treatment and control. They work well when you have enough comparable regions and your channel can be targeted geographically. The main hazard is spillover — media leaking across boundaries — and pre-existing differences between regions that have nothing to do with your test.
Audience holdouts withhold advertising from a randomly selected user segment within the platform. Cleaner randomization, but only available where the platform supports it, and you are partly trusting the platform's own measurement of a test that evaluates that platform.
Run long enough to cover your purchase cycle. A two-week test on a product with a six-week consideration window measures the beginning of the effect and calls it the whole effect. The test must outlast the lag between exposure and purchase, plus enough time to accumulate signal.
Size for the effect you are looking for. The smaller the true lift, the more data you need to distinguish it from noise. Teams routinely run tests underpowered to detect the effect sizes they care about, then read the inconclusive result as "no effect." Those are very different findings. Work out before launching how large a lift the test can actually detect.
Expect the number to be lower than your platform reports. Platform-reported conversions include people who would have bought anyway. That gap is the entire point of running the test, and it is not a malfunction when it appears.
The dependency all three share
Here is the part the methodology debate skips.
MMM ingests conversion counts over time. Incrementality compares conversion counts between groups. MTA reconstructs journeys ending in conversions. Every one of them takes your conversion event stream as ground truth. None of them can tell you when that stream is wrong, because none of them has an independent reference to check it against.
Consider a holdout test where your event capture is lossy — say browser-side pixel tracking missing a meaningful share of conversions, with the loss concentrated in Safari and iOS. Your treatment group is being advertised to, so it likely skews slightly toward mobile and social placements, where measurement loss is worst. The undercounting hits treatment harder than control. Your measured lift comes out understated, and you conclude a channel that is working is not. There is nothing in the test output that reveals this. The confidence interval is computed on the numbers you fed it.
The failure modes that matter:
- Missing conversions bias any comparison where the loss rate differs between the groups being compared — which it usually does, because the groups differ in exactly the browser and device mix that determines loss rate.
- Duplicate conversions, from a pixel and a server event both firing without proper deduplication, inflate counts unevenly across channels and time periods.
- Timestamp drift — recording conversion time as processing time rather than actual event time — smears conversions across time buckets. MMM in particular is sensitive to this, since its whole method is fitting temporal relationships between spend and outcome.
- Inconsistent conversion definitions across sources, so the "conversions" in your MMM input are not the same population as the ones in your holdout readout.
None of this is an argument that measurement methodology does not matter. It matters a great deal. The argument is narrower: the sophistication of your measurement method sets a ceiling on your insight, and the quality of your event data sets the floor. Investing heavily in the former while ignoring the latter produces a rigorous analysis of the wrong numbers.
Where to start
If you are running none of these formally, the order that tends to pay off fastest:
- Audit the event stream first. Reconcile platform-reported conversions against your order table for a recent period, split by browser and device. If those disagree substantially, fix that before investing in any measurement methodology — everything downstream inherits the discrepancy. Complete server-side capture is usually where the gap closes.
- Run one incrementality test on your largest channel. One well-designed holdout on your biggest line item teaches more than a year of dashboard-watching.
- Add MMM once you have the history — realistically two or more years of reasonably clean weekly data, plus the analytical capacity to interpret it.
- Keep MTA for within-channel tactics, and stop using it to arbitrate between channels.
The teams that measure well are rarely the ones that picked the right framework. They are the ones who know what their data is missing, and size their confidence accordingly.


