Blog

Your Marketing Mix Model Might Be Lying, and the Vendors Wrote It Down in the Docs

Geometric vector illustration of a faceted model surface lifting away from the terrain it is meant to describe

An AI marketing mix model does not replace incrementality testing, and you do not need me to argue that, because Google and Meta already put it in their own documentation. Meridian's calibration page says incrementality experiments are "perhaps the strongest basis" for setting a channel's prior. Robyn scores the gap between its own predictions and experimental lift results as a third objective function, and calls randomized controlled trials the gold standard. The accuracy question for marketing mix modeling ai turns on validation rather than on math. The math is good, and better than it was two years ago. Almost nobody checks the output against an experiment.

The vendors are not hiding this, which is stranger than if they were

A marketing mix model, or MMM, is a regression fitted to aggregate time series: weekly spend by channel, weekly sales, plus controls for price, promotion and seasonality. Incrementality is the causal difference between what happened and what would have happened without the spend. A prior is what you tell a Bayesian model to believe before it sees any data, and calibration is the act of building that prior out of experiment results instead of out of guesses.

Meridian's documentation page on calibration, last updated 2026-09-01, defines calibration as "the process of using experiment results and other domain knowledge to set channel-specific priors." Every Meridian doc page carries a banner reading "Meridian GeoX is now available. Run geo experiments to calibrate your MMM." Meta's Robyn implements calibration as an objective function named MAPE.LIFT, sitting alongside prediction error and business error, and takes GeoLift and Conversion Lift results as the input.

Then Meridian goes further than I expected. Its model health report produces a channel calibration score and flags any channel scoring under 67.5, assembled from four checks: implausibly high ROI, implausibly low ROI, high variance ROI, and potential bias from missing confounders. The page explains why the feature exists, and the explanation is a concession: "Because running experiments across all media channels is logistically complex and expensive, Meridian helps identify channels where the model exhibits high uncertainty, potential bias, or unrealistic baseline estimates to prioritize for testing." Google shipped a feature whose job is telling you which of Google's own model outputs to distrust.

The claim that AI MMM replaces incrementality testing is not coming from the documentation, then. It survives in decks, and in dashboards, where a response curve renders as one smooth confident line and nothing on screen marks which parts of it were ever measured.

Why two models that fit your data equally well recommend different budgets

David Chan and Michael Perry published "Challenges and Opportunities in Media Mix Modeling" at Google in April 2017, and it is still the clearest statement of the problem. Three years of national weekly data is 156 data points. From those you are asked to model 20 or more channels, at 3 or 4 parameters each for lag and diminishing returns, against a rule of thumb of 7 to 10 data points per parameter. Typical MMMs fall short of that, and the paper says so plainly.

What you get is ambiguity rather than noise. Advertisers spend across channels in correlated ways, because that is what sensible media planning looks like, and correlated inputs produce coefficient estimates with high variance and, in the paper's words, "bad attribution of sales to the ad channel." Chan and Perry fitted five plausible models to one real US weekly dataset. Every one hit an R2 of 0.98 or 0.99, four of the five hit out of sample MAPE of 6 to 8 percent on the last 12 weeks, and they still differed in their sales predictions by up to 50 percent and disagreed about how to allocate budget, including about the single largest channel the advertiser had.

Ryan Dew at Wharton, Nicolas Padilla at London Business School and Anya Shchetkina at Wharton put a sharper version on arXiv in August 2024, titled "Your MMM is Broken." They ran 2,187 simulation settings at 100 datasets each, and found that a nonlinear reading and a time varying reading of the same data conflate routinely, with noise in the response as the largest single driver. On one simulated dataset where both fit equally well, one implied optimal spend of $12,072.60 and the other implied $7,451.50, a gap close to the entire range of the training data. Both readings were defensible.

The number that should bother you came from the run with perfect controls

On 21 August 2026, three weeks ago, a preprint went up on arXiv under the identifier 2608.21128, by Niklas Heusch, whose contact address resolves to a zalando.de domain. It builds a synthetic online retailer running paid search, social and television across 156 weeks, with endogenous budget allocation and performance chasing bidding written into the generating process, then fits a mix model specified the way a competent practitioner would specify one: a promotion dummy, an observed promotional price level, and three annual Fourier harmonics for seasonality, the standard seasonal controls in Robyn, Meridian and pymc-marketing.

True paid search return in that world was 4.20x. The model reported 10.61x, with a 90 percent credible interval of 6.56 to 14.36 that did not contain the truth.

Then the paper runs it again with oracle controls, meaning the model gets handed information no real analyst could have. It reported 8.41x, interval 6.91 to 9.81, and still missed. That second run is worth sitting with, because it suggests the gap is not a data quality problem that a richer dataset closes.

The parameter that moved most was carryover. True adstock was 0.20 and the model estimated 0.50. A channel where half of last week's spend carries into this week is one you can pulse, and flight, and rest between bursts, and a channel with 20 percent weekly carryover will not survive that treatment. Same spreadsheet, different media plan.

I want to be careful about what this is. It is a simulation, single author, not peer reviewed, and the comparison is only possible at all because somebody wrote the true answer down first. The factor of roughly 2.5 is a property of that simulation rather than a constant to expect in your account, and the exercise inherits its own assumption that reality is exactly geometric adstock and logistic saturation. It demonstrates a mechanism, not a magnitude.

Paid search is the worst channel for this, and the correction is at the query level

In July 2018 a team of Google researchers including Aiyou Chen, David Chan and Mike Perry published "Bias Correction For Paid Search In Media Mix Modeling." Its abstract reports the most alarming case study number in this literature: a naive MMM estimate of paid search ROAS coming out almost 15 times larger than the experimental result, with a simple category search volume adjustment shrinking the gap and leaving the estimate still around seven times too large. I am citing that abstract rather than the case study tables, which I have not read, and in a piece arguing that people should check numbers before repeating them it would be poor form not to say so.

The reason is targeting. Search ads get served to people who already typed a query indicating interest, so the spend follows demand that would partly have converted anyway, and a regression on weekly totals has no way to pull those apart. The correction the paper derives is to use relevant search query volume as a demand control variable.

That is a query level fix for an aggregate level failure, and it is the part you can act on today without commissioning anything. A mix model sees one number per week for a channel called Search. The query report sees whether that number was branded navigational traffic arriving regardless, or genuinely new nonbrand demand, and the mix moves week to week while bidding algorithms chase performance signals. Two accounts with identical weekly spend curves can carry completely different incrementality, and nothing in the aggregate series tells them apart. Keeping that composition visible before it gets averaged away is most of what our paid search management work actually is, and it is why our measurement and conversion tracking setups start at query and campaign grain instead of the channel total.

Meridian's own potential bias check computes the maximum absolute correlation between a channel's execution and any control variable across geos, which is sound at the level it operates on, and still blind to composition inside the channel.

Where MMM is genuinely good, and what a geo test actually costs

Credit where it belongs. Chan and Perry's argument for why advertisers reach for MMM at all still holds: to answer what happens if you move budget from television to social next year, you would need many experiments under many conditions, which for most advertisers is infeasible. MMM survives privacy changes that break user level measurement, and it covers offline channels no pixel reaches, and it hands you a response curve rather than a single point. Meridian has moved fast on top of that, adding non media variables in September 2025, a Scenario Planner in February 2026, and GeoX in May 2026.

Geo-lift is no free oracle either. A geo-lift test splits markets into treatment and control, cuts or zeroes spend in the treatment markets, and measures divergence against a synthetic control built from the untreated ones. Meta's own GeoLift documentation admits the method "is subject to biases due to inexact matching," and warns that "running a geo-test without a robust prior power analysis leads to a high chance of failing to find lift, even if it actually happened." An underpowered null handed to a CFO does more damage than a biased model does.

Cost is real as well. Google cut its minimum incrementality experiment budget from near $100,000 to $5,000 on 11 November 2025, which genuinely opened testing to mid market advertisers, and also reported that most advertisers run one to two incrementality studies a year. The Heusch design needs four tests at different spend levels, 16 weeks each including cooldown, to pin down saturation on one channel. Nobody is running that.

The honest position lands somewhere short of testing everything. An uncalibrated model is an unvalidated model, every vendor quoted in this piece says so in writing, and one test a year aimed at your largest or least plausible channel beats none.

FAQ

Is AI marketing mix modeling accurate? It is precise more often than it is accurate. Chan and Perry's 2017 case showed five models all scoring R2 of 0.98 or 0.99 while disagreeing about sales by up to 50 percent. Fit and correct causal attribution are separate properties, and only the second one decides budget well.

Does MMM replace incrementality testing? Not according to the vendors. Meridian's docs call incrementality experiments "perhaps the strongest basis" for setting a prior, ship a score flagging channels below 67.5 for testing, and run a site wide banner telling you to calibrate with geo experiments.

How many incrementality tests do I actually need? Google reports most advertisers run one to two a year, and no published threshold makes a model safe. Test the channel your model treats as implausible or uncertain first, which is what Meridian's calibration score identifies. Its default ROI prior is LogNormal(0.2, 0.9), an 80th percentile range of 0.11 to 2.64, which is a very wide belief to start from.

Is geo-lift ground truth? Closer than a mix model, and not perfect. Meta's GeoLift docs concede inexact matching bias, and geographic spillover between treated and control markets biases the measured difference downward.

What I would actually do with this

Look at whether search query volume appears anywhere in your model as a control variable. If it does not, the model has no mechanism for separating spend that created demand from spend that followed it, and the channel most exposed to that is probably the one you spend the most on.

The vendors documented all of this in public. The gap between what their docs say and what gets presented in a quarterly review is the part worth closing.

All insights and articles

Reading about it is the easy part

Stop guessing. Start knowing your growth potential.

Get a free, no obligation audit of your paid search account, real math from your own query data, not a sales pitch.

The Optimal Path to Conversion

With precision bidding strategy, delivering the right Keyword/Match Type, to the right Ad Message, to the right Landing Page.

Wasteful Themes Pruned - Negatives
See how queryDNA works →