A paper posted to arXiv on August 21, 2026 by Niklas Heusch shows a standard marketing mix model reporting 10.61x return on ad spend for a paid search channel whose true return was 4.20x, and demonstrates that geo-experiment time-series can recover the correct figure.

The number that will travel is 10.61 against 4.20. It comes from a simulation in which every parameter is known, which is the only setting where such a comparison is possible at all, and it describes a model specified the way a competent practitioner would specify one: a promotion dummy, an observed promotional price level, three annual Fourier harmonics for seasonality. Those are the standard seasonal controls in Robyn, Meridian and pymc-marketing. They were not enough. The model overstated the channel's return by a factor of roughly 2.5, and its 90 percent credible interval, running from 6.56 to 14.36, did not contain the truth.

The second number is the one that matters more. When the same model was handed the actual confounders from the data-generating process - the true promotion state, the true price level, the latent seasonal component, product-quality drift, market sentiment - it still returned 8.41x, with an interval of 6.91 to 9.81 that also excluded 4.20. No practitioner possesses those covariates. The specification exists in the paper as a diagnostic rather than a proposal. Its purpose is to bound what better controls could ever achieve, and the bound is roughly double the true return.

That is a different claim from the familiar one about data quality. It says the gap is not a measurement problem that a richer dataset closes.

Who wrote it and where it sits

The paper is titled "Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments" and carries the identifier arXiv:2608.21128v1, filed under Applications in the statistics section, with a submission stamp of 21 August 2026. It lists a single author, Niklas Heusch, and no institutional affiliation line. The contact address at the head of the paper resolves to a zalando.de domain, which is the only signal in the document about where the work sits commercially. A companion paper, cited as Heusch 2026, is credited with generating the synthetic dataset used throughout, and the paper states that the notebook producing that dataset is public.

It arrives into a market that has spent two years rebuilding aggregate measurement. Google opened Meridian globally in January 2025 after testing with hundreds of brands, added non-media variables, channel-level contribution priors and binomial adstock decay in September 2025, launched a Scenario Planner in February 2026, and announced Meridian GeoX and Meridian Studio on May 5, 2026. GeoX is a geographic incrementality testing tool whose stated function is to feed experiment results into the mix model. The Heusch paper is, in effect, an argument that the standard way of performing that handover throws away most of the information the experiment produced.

What standard practice does with an experiment

A geo-experiment divides a market into geographic units, typically Designated Market Areas or metropolitan regions, and assigns them at random to treatment or control. During the test window, spend in the treatment group is modified, often reduced to zero in what the industry calls a go-dark test, while the control group carries on spending normally. A well-designed test has three phases: a pre-test period establishing that both groups follow parallel trends, a test period in which spend differs, and a cooldown period in which spend returns to normal and residual effects decay.

The raw output is rich. Both groups produce complete time series of outcomes and spend, typically at daily granularity, across all three phases.

Standard practice then collapses all of it into one number. Total incremental revenue over the test period is divided by total incremental spend over the test period, producing a point estimate of return on ad spend. That figure answers a narrow question about average return at the spending level tested. The paper's objection is stated plainly: summing over the test period and dividing totals discards every piece of information about dynamics, meaning how effects evolve across weeks and how they respond to different spending levels.

Two specific signals disappear in that summation, and both are visible in the raw series.

The first is carryover. When spend stops in the treatment group at the start of a go-dark test, the outcome gap between the groups does not open instantly. It widens gradually as the treatment group's accumulated advertising effect depletes while the control group's is maintained. When spend resumes at the start of the cooldown, the gap does not close instantly either. It narrows gradually as the effect rebuilds. The rate of those two transitions identifies the decay parameter directly. Fast transitions, where the gap opens or closes inside a week or two, imply low carryover. Slow transitions, where convergence takes many weeks, imply high carryover.

The second is diminishing returns. Tests conducted at different spending levels trace out the response curve. If doubling spend doubles the outcome gap, returns are proportional. If doubling spend yields less than double the gap, saturation is present and its curvature can be estimated.

A single ROAS figure contains neither signal. The paper's framing of its own contribution is a shift from "experiments for measurement" to "experiments for structural estimation".

Three parameters, and what each one governs

Marketing mix models translate spend into incremental sales through two transformations, and the paper is specific about the arithmetic of both.

Adstock captures persistence. The adstock in a given week is a weighted sum of spend in that week and in the preceding weeks inside a lookback window, with each past week weighted by the decay rate raised to the number of weeks elapsed. The denominator normalises the weights to sum to one, so that a thousand euros of spend produces a thousand euros of total adstock distributed across weeks rather than concentrated in one. When the decay rate is zero, only the current week matters. At 0.7, spending from several weeks earlier still contributes substantially.

Saturation captures diminishing returns to that adstock. The paper uses the logistic saturation function, which rises steeply at low adstock and flattens at high adstock. A higher curvature parameter means diminishing returns set in earlier, at lower levels of accumulated advertising. An effectiveness coefficient then converts the saturated adstock into euros of incremental revenue.

Three parameters therefore govern how spend on any given channel translates into sales: persistence, speed of diminishing returns, and overall effectiveness. Every budget allocation decision an MMM produces is a function of all three. A May 28, 2026 developer episode walked through the equivalent machinery inside Meridian, where Hill curves handle saturation and adstock handles carryover.

Why the observational model cannot recover them

The identification problem is not subtle, and the paper states it in one line: marketing budgets are not randomly assigned.

Companies increase advertising ahead of periods when demand is expected to be high. Automated bidding systems raise spend when conversion rates rise, and high conversion rates frequently reflect strong underlying demand rather than advertising effectiveness. Budgets respond to last quarter's results, which embed last quarter's shocks. Strategic decisions coordinate campaigns with promotional calendars. Each of these creates correlation between spend and the error term, and the model attributes that correlation to causation.

For a mix model to yield causal estimates, the error term must be independent of marketing spend conditional on the observed controls. The paper's appendix works through why that condition is virtually impossible to satisfy in practice. Two mechanisms in the synthetic data survive even perfect controls, and the paper names both. The bidding algorithm responds to realised weekly performance, including its random component, which no covariate an analyst could hold constant can absorb. And the true baseline combines its components multiplicatively while the controls enter the regression linearly, leaving residual variation that the endogenous spend tracks.

That is why the oracle specification matters. Described in the paper as not a feasible model but a diagnostic, it bounds what any improvement in controls could achieve. It cut the bias from 10.61x to 8.41x and stopped there. The remaining distance to 4.20x cannot be closed by finding better data, because the model was already handed the best data that exists.

Behavioural and algorithmic feedback of exactly this kind is now standard in paid search. Algorithmic bidding systems set prices at auction time from more than twenty signals, adjusting continuously to observed performance. The endogeneity the paper simulates is not a stylised nuisance. It is the default operating condition of the channel being measured.

The differencing equation

The mechanism the paper proposes is simple once stated. Treatment and control groups can be treated as two parallel universes, identical in every respect except for spending on the tested channel.

Outcomes in each group are written as a common component plus the contribution of the tested channel plus idiosyncratic noise. The common component contains everything else: baseline demand, contributions from all other channels, pricing, promotions, seasonality, and every unobserved factor. Because randomisation ensures the two groups differ only in spend on the tested channel, that common component is identical in both.

Subtracting one from the other eliminates it. The intercept cancels. Other channels cancel, because their spend is identical in both groups. Observed control variables cancel, because their values are identical. Unobserved confounders cancel, because randomisation balanced them. What remains is an equation containing only the three parameters of the tested channel and observable quantities: the outcome difference on the left, spend in both groups on the right.

No conditional independence assumption is required. The paper is explicit that its method does not need controls to capture all confounders and does not need a model of how budgets are set. Identification comes from the experimental design rather than from hoping the covariates were sufficient.

The paper situates the estimator against three familiar econometric frames. Like difference-in-differences, it compares treatment and control over time, but it estimates structural parameters rather than average treatment effects and uses the full time series rather than pre and post comparisons. Randomisation provides an instrument for spending, but the approach estimates the structural form directly rather than through two-stage least squares. And the parallel universes framing maps onto the Rubin causal model, with treatment and control as draws from the same potential outcome distribution differing only in realised treatment.

Data preparation and pooling

Each observation requires outcomes in both groups at a given time, plus spend in both groups at that time and across the preceding weeks of the lookback window, since those determine the adstock.

The paper reshapes geo-test data into wide format, where each row holds the outcome difference and the current plus lagged spend for both groups. The property that matters is self-containment: every row carries all the lagged values it needs, so no cross-row temporal dependencies remain.

That makes pooling across experiments mechanical. Rows from different tests are stacked. Observations from a test run in January can sit alongside observations from a test run in June, each providing independent information about the same underlying parameters. The Bayesian prior approach used by existing calibration methods cannot distinguish experiments conducted at different spending levels without ad-hoc weighting; stacked rows carry their own spending levels with them.

Scaling is handled explicitly, because the parameters are not scale-invariant. A test run on a subset of a country produces smaller absolute figures than the country-level model it is meant to calibrate. The paper divides treatment outcomes and spend by the treatment group's share of the country and control figures by the control group's share, transforming the data to what would have been observed had each group covered the whole market. Any transformation the mix model applies for numerical stability, such as dividing by maximum values, is then applied to the geo-test data as well.

The estimation setup

Estimation runs in PyMC using the No-U-Turn Sampler with four chains, 1,000 tuning iterations and 2,000 sampling iterations each, at a target acceptance rate of 0.95. The observational benchmarks use 1,000 sampling iterations. Convergence is assessed at R-hat below 1.01 and effective sample size above 400, and the paper reports that all runs pass. Estimation completes in under 30 seconds on a modern laptop.

The priors are the package defaults of pymc-marketing version 0.19: a Beta(1, 3) prior on adstock decay reflecting limited carryover for digital channels, a Gamma(3, 1) prior on saturation with a mode at 2, a half-normal prior with scale 2 on effectiveness, and a half-normal prior with scale 0.5 on observation noise.

The design decision worth noting is that the observational benchmark uses the same media-parameter priors as the structural model. Its remaining priors are normal with scale 0.3 on control coefficients, normal with scale 0.2 on Fourier coefficients, normal centred on 0.5 with scale 0.2 on the intercept, and half-normal with scale 0.1 on the noise scale. The oracle specification's control coefficients carry a wider normal prior with scale 2, which the paper explains as removing any argument that the priors held the perfect covariates back. The comparison between methods therefore cannot be attributed to prior choice.

Inside the probabilistic model

The paper publishes the model structure in plate notation, generated directly from PyMC's graph function, and the arrangement of the nodes carries the argument.

Observed inputs sit at the bottom of the hierarchy as rectangular nodes: marketing spend for the control group and for the treatment group, each containing the lookback window of current plus lagged spend. The dimensions are stated as 16 weeks by six lags, matching the test length and the lookback.

The three structural parameters sit at the top as elliptical nodes, and they are shared between both groups. That sharing encodes a substantive assumption rather than a computational convenience: the marketing channel is assumed to work the same way regardless of which group a region was assigned to. Randomisation is what makes the assumption defensible.

Between the two levels sit the parallel transformations, computed separately for each group and identical in form: adstock, then saturation, then impact. The predicted outcome difference is the difference between the two impact terms, and the observed outcome differences are modelled as normally distributed around it.

The detail that matters is positional. The plate, meaning the rectangle drawn around the observation-level nodes, encloses the transformations because they are computed for every time period. The parameters sit outside it. They are estimated once and applied identically to both groups across every observation, which is what allows rows from a January test and a June test to inform the same three numbers.

The synthetic data and four experiments

The demonstration uses synthetic data simulating an online retailer with three marketing channels - paid search, social media and television - across 156 weeks. The generating process incorporates realistic seasonality and trend, endogenous budget allocation in which spend responds to expected demand, algorithmic bidding that chases performance signals, and known adstock and saturation effects.

Every parameter value, magnitude and business mechanism is fictional, chosen to sit inside the range of published industry benchmarks. The complete data-generating process, including the true value of every parameter, is public. That is the whole point of the exercise: ground truth exists only in simulation.

For the paid search channel, referred to throughout as PLA, the true parameters are an adstock decay of 0.2, meaning 20 percent weekly carryover; a saturation parameter of 0.008 per thousand euros of adstocked spend, implying half-saturation at roughly 137,000 euros of weekly adstocked spend; an effectiveness coefficient of 1,100 in thousands of euros per unit of saturated adstock; and a true return on ad spend of 4.20x. Because estimation divides each series by its maximum, the parameters as reported in the tables live on that scale, where the true saturation figure is approximately 1.71 and the true effectiveness figure approximately 0.18. The scaling constants are a maximum weekly paid search spend of 214,000 euros and maximum weekly sales of 6.03 million euros.

Four geo-experiments are generated from that data, deliberately placed at different points in the three-year window and at different spending levels:

  • Test 1, weeks 20 to 23, at a mean control-group paid search spend of 54,000 euros per week
  • Test 2, weeks 55 to 58, at 129,000 euros
  • Test 3, weeks 100 to 103, at 64,000 euros
  • Test 4, weeks 140 to 143, at 57,000 euros

Each test includes four weeks of pre-test period, four weeks of test period with paid search spend set to zero in the treatment group, and eight weeks of cooldown. All 16 weeks of every test enter the estimation, with lagged spend drawn from the full history, so that the gap-opening and gap-closing transitions contribute to the likelihood in full.

The spread of spending levels is not incidental. Jointly, the four tests probe adstocked weekly spend from roughly 43,000 to 194,000 euros, which is close to the full range observed across the three years. That variation is what identifies the saturation curvature. Both experimental groups are simulated directly at country scale, which removes the rescaling step without changing anything else about the exercise.

In the first test, treatment group sales fall relative to control during the test window, and the gap closes almost immediately once spending resumes - a pattern the paper reads as consistent with the low true carryover of 0.2.

The results table

The comparison covers four methods against a known truth of 4.20x.

MethodROAS90% credible intervalContains truth
True4.20x--
MMM, realistic controls10.61x6.56 to 14.36No
MMM, oracle controls8.41x6.91 to 9.81No
Structural, 2 tests4.31x3.87 to 4.75Yes
Structural, 4 tests4.14x3.79 to 4.48Yes

Two structural estimates bracket the truth from either side and both intervals contain it. Precision improves as more experiments are pooled, with the four-test interval running 0.69 wide against 0.88 for two tests, but the two-test version already produces an accurate point estimate. The practical reading is that pooling helps and that a firm running two well-designed tests on a channel is not obviously worse off than a firm running four.

The observational estimates are wrong in the same direction, which is the direction that costs money. A channel reported at 10.61x when it actually returns 4.20x invites more budget, and the marginal euro spent on that assumption is being priced against a response curve that does not exist.

Beyond ROAS: the parameters themselves

Return on ad spend is a summary statistic. The underlying parameters are what a planner uses to decide how much to spend and when.

MethodAdstockSaturationEffectiveness
True0.201.710.18
Structural, 2 tests0.18 (0.08 to 0.29)2.27 (0.89 to 3.47)0.17 (0.11 to 0.33)
Structural, 4 tests0.19 (0.09 to 0.28)1.96 (0.67 to 3.13)0.20 (0.11 to 0.43)
MMM, realistic0.50 (0.34 to 0.63)1.28 (0.40 to 2.63)0.82 (0.28 to 1.87)

The structural estimates place the true value inside every 90 percent credible interval. The adstock posterior centres almost exactly on the truth, at 0.19 against 0.20, and the paper attributes that precision to the pre-period and transition weeks of each test, which enter the likelihood in full rather than being summarised away.

The observational model's adstock estimate of 0.50 against a true 0.20 is the failure with the clearest operational consequence. A model believing that half of last week's spending carries into this week describes a channel that can be pulsed, flighted and rested. A channel with 20 percent weekly carryover cannot. Those are different media plans, and the effectiveness coefficient - estimated at several times the truth - compounds the error into the inflated return figure.

The saturation and effectiveness posteriors from structural estimation are individually wider, and the paper explains why rather than glossing it. The two parameters trade off along a ridge, with a posterior correlation of -0.68, because a flatter response curve paired with a larger coefficient produces nearly the same predicted response as a steeper curve paired with a smaller one, across the range of spending the experiments actually covered. What the data pin down tightly is the joint implication of that ridge: the response curve over the tested range, and the return it implies. That is why the ROAS intervals are considerably narrower than the marginal parameter intervals.

Adstock, by contrast, is identified separately, by the gap-opening and gap-closing transitions rather than by the level of the effect.

What the experiments do not pin down

The paper is unusually direct about the boundary of its own claims.

Over the range of adstocked spending the four tests probed, the posterior response curves lie on top of the true curve. Beyond that range they fan out. The paper describes this as an honest statement of what the experiments do and do not establish, and draws the operational consequence: recommendations far outside the tested range, such as a large budget expansion, rest on the parametric form of the saturation function rather than on experimental variation.

That distinction is easy to lose in a dashboard. A response curve rendered as a smooth line from zero to some upper bound looks equally authoritative along its whole length. The evidence supporting it is not evenly distributed.

Four further limitations are named, and the paper notes that most are shared with mix modelling generally rather than specific to the method.

Functional form assumptions come first. The approach assumes geometric adstock and logistic saturation correctly describe the true process. The paper argues this is an advantage rather than a liability, because differencing removes confounders and leaves an equation containing only the parameters of interest, so alternative specifications - geometric against delayed adstock, logistic against Hill saturation - can be compared using standard model selection criteria. In observational estimation, confounding makes it difficult to distinguish misspecification from omitted variable bias.

Geographic spillovers come second. The method assumes advertising in control areas does not reach treatment consumers. Where it does, for example when consumers travel between regions, the outcome difference understates the true effect. This constrains all geo-experiments equally.

Parameter stability comes third. Parameters are assumed constant across experimental periods. Where effectiveness varies substantially over time through creative fatigue, competitive dynamics or shifting preferences, the estimates reflect an average across the experimental windows. The paper cites Dew et al. (2024) as the reference for time-varying alternatives, which is the second paper considered here.

Experimental requirements come fourth, and this is the one with a direct operational cost. The method needs tests with enough temporal variation to identify dynamics. A two-week test may not generate enough adstock variation to pin down a decay rate. Tests must span different spending levels to identify saturation curvature. The paper acknowledges that these requirements may exceed what standard geo-test designs provide, and argues the informational gains justify enhanced protocols.

Test length and spending variation are precisely the dimensions that cost money in a live account. Four weeks dark on a performance channel across a meaningful share of a market is a real revenue line, and a design that requires four such tests at different spending levels is a materially larger commitment than a single lift study. Google cut the minimum experiment budget for incrementality testing from around $100,000 to $5,000 on November 11, 2025, and reported that most advertisers conduct roughly one to two incrementality studies a year. One to two studies a year is below what the structural method needs to identify saturation.

How it compares with what the tools already do

Existing calibration approaches all reduce experiments to aggregate statistics, and the paper works through three by name.

Google Meridian, citing Zhang et al. 2024, reparameterises the model to include return on ad spend as a direct parameter, so that experimental point estimates can serve as informative Bayesian priors. The paper describes this as an elegant way to fold aggregate experimental evidence into estimation, and identifies the limitation: it cannot distinguish between experiments conducted at different spending levels or time horizons without ad-hoc weighting. It identifies the level of channel effectiveness but not how effectiveness varies with spend or persists over time.

Meta Robyn, citing Runge et al. 2023, treats calibration as a third objective in multi-objective optimisation, penalising deviation from experimental lift estimates alongside fit and allocation criteria. This steers model selection toward experimentally consistent regions of the parameter space. Without modelling the temporal structure of experimental outcomes, the method constrains the aggregate effect without informing its decomposition into adstock and saturation.

pymc-marketing, citing Orduz 2024, adds lift test observations directly to the model likelihood, treating experimental estimates as points on the saturation curve. Where experiments at different spending levels are available, this can inform saturation. Because it uses aggregate lift rather than time-series outcomes, it cannot identify how long effects persist after spending stops.

The paper's own contribution uses three distinct features of the data: the rate at which the outcome gap emerges and closes, to identify adstock decay; the relationship between spending differences and outcome differences, to identify saturation curvature; and the level of outcome differences, to identify effectiveness. All three emerge from a single estimation rather than from separate heuristics.

Three integration paths into a full multi-channel model are set out. Informative priors take the posterior from structural estimation as the prior for the corresponding channel, which requires specifying a joint prior capturing the correlations between parameters - a non-trivial requirement given the documented ridge between saturation and effectiveness. Joint likelihood incorporates the differencing equation directly into the mix model likelihood, treating observational series and geo-test observations as arising from the same parameters; the paper notes that the existing pymc-marketing codebase provides a natural foundation for this, since it already incorporates lift test observations. Two-stage estimation computes the channel's contribution from structural estimates, subtracts it from total sales, and estimates the remaining model on residualised outcomes; this is computationally simpler but does not propagate uncertainty from the first stage.

Robustness, and the bias that does not move

The paper re-runs the structural estimation under two alternative prior sets.

Prior setAdstock priorSaturation priorROAS, 4 tests
BaselineBeta(1, 3)Gamma(3, 1)4.14 (3.79 to 4.48)
DiffuseBeta(1, 1)Gamma(1, 0.5)4.06 (3.72 to 4.41)
InformativeBeta(2, 8)Gamma(5, 2)4.14 (3.82 to 4.45)

Point estimates shift by less than 2 percent and every interval contains the true value. The same test on the other side of the comparison produces the more telling result. Re-estimating the realistic observational model under an alternative prior set - Beta(2, 1) on carryover, Gamma(2, 1) on saturation, normal with scale 0.5 on the media coefficients - moved its estimate from 10.61x to 11.38x. The bias is driven by the likelihood rather than by the priors, and loosening or tightening prior beliefs does not touch it.

That result closes off the most common defence of a mix model producing an implausible number. Adjusting the priors is the standard remedy when an MMM returns a figure the commercial team does not believe. In this simulation it made the figure worse.

Every result here is synthetic

The paper's empirical section runs entirely on simulated data, and states so. There is no validation against a real advertiser's sales, no holdout on live campaign data, and no comparison with an audited outcome.

This is a genuine constraint rather than an oversight, and the reason is structural. Ground truth for advertising effectiveness does not exist in observational data. The only setting in which an estimator can be checked against a known true return is one where the analyst wrote the true return in the first place. The cost is that the demonstration inherits every assumption embedded in the simulation, including that the true process is exactly the geometric adstock and logistic saturation form the estimator assumes.

The paper's own limitations section concedes the functional form point. It does not extend the concession to the possibility that the observational model's failure is partly an artefact of a data-generating process built to include endogenous allocation and performance-chasing bidding. Those mechanisms are real, and the paper's citations to Gordon et al. 2019 and Gordon et al. 2023 provide independent empirical support that observational methods fail to recover experimental effects even with hundreds of millions of observations. But the magnitude of the failure reported here - a factor of 2.5 - is a property of this simulation rather than a general constant.

The practitioner reaction

The paper circulated on LinkedIn accompanied by a chart lifted from its results section, showing the five return figures with their credible intervals: ground truth at 4.2x, the realistic model at 10.6x, the oracle at 8.4x, and the two structural estimates at 4.3x and 4.1x.

The note attached to the share was not an endorsement of a vendor or a methodology so much as a pre-emptive fence around the comment thread. It anticipated that mix modelling vendors would have thoughts about why their products deliver superior results to this implementation, stated that such comments were not invited, said the paper had been enjoyable to read and that this was the reason for sharing it, and warned that anything read as promotional would be deleted.

That is a small piece of social media housekeeping, and it is also an accurate description of the commercial environment the paper lands in. A finding that a standard mix model overstates return by 2.5 times is, for every vendor selling a mix model, either an attack or a marketing opportunity depending on whether the vendor believes its own implementation differs. The paper does not test any commercial product. It tests a specification that the paper describes as what a practitioner would actually run, using the standard seasonal controls of three named open-source frameworks. The distinction between "the specification most people use" and "this vendor's product" is real, and it is also the distinction that will be blurred fastest in the comment threads.

The post recorded 67 reactions, eight comments and two reposts at the point of capture, which is a modest circulation for a technical paper and a reasonable one for a preprint on Bayesian estimation.

The second paper: a prior finding the first one depends on

The Heusch paper cites Dew et al. (2024) as its reference for the identification failure it is trying to route around. That paper, "Your MMM is Broken: Identification of Nonlinear and Time-varying Effects in Marketing Mix Models", was posted to arXiv on 14 August 2024 under the identifier 2408.07678v1 in the econometrics section, and carries a title-page date of August 15, 2024. The one-day gap between the submission stamp and the stated date is the sort of detail that appears in most preprints and is noted here only for the record.

Its authors are Ryan Dew, an assistant professor of marketing and Govil Family Faculty Scholar at the Wharton School of the University of Pennsylvania; Nicolas Padilla, an assistant professor of marketing at London Business School; and Anya Shchetkina, a doctoral student at Wharton. The paper states that the authors contributed equally and are listed alphabetically.

The two papers describe different failures, and understanding how they interact matters more than either does alone. Heusch attacks endogeneity: the model cannot separate the effect of advertising from the demand that caused the advertising. Dew, Padilla and Shchetkina attack something that survives even when endogeneity is absent. Their central result holds when spending is exogenous, meaning independent of the error term.

Two model families that cannot be told apart

Modern mix models capture diminishing returns through nonlinear response functions. They also, either explicitly or implicitly, allow marketing effectiveness to change over time. The paper's claim is that these two effects are frequently not separately identifiable from standard mix model data, and that they imply fundamentally different optimal allocations.

The illustrative example in the introduction is a scatter of revenue against advertising spend that looks unmistakably nonlinear, with a clear inflection and apparent diminishing returns. A flexible nonlinear model captures most of the variation. The true generating process is a linear model in which the coefficient varies smoothly over time in a cyclical pattern, with spending set to follow that cycle with a lag.

The theoretical result runs in two directions and the conditions differ.

A time-varying linear model can approximate a nonlinear one when the nonlinearity is sufficiently smooth. The paper works through the algebra using a Taylor expansion around mean spend, showing that a heavily regularised time-varying model reduces to ordinary least squares, whose coefficient equals the derivative of the true response function at mean spend plus two error terms, one of which depends on the second derivative. The approximation improves as the function behaves more smoothly. It also improves when spending itself evolves slowly, because a smoothly varying coefficient cannot track a nonlinear response when spend jumps around between periods. In practice spending does evolve slowly: firms set budgets in an autocorrelated way, increasing by a percentage rather than resetting from scratch, and stock variables such as adstock impose smoothness on the input series by construction.

The reverse direction is narrower. A static nonlinear model can approximate a time-varying linear one only when the time-varying effectiveness can be written as a function of spending. The paper proves this condition is both necessary and sufficient. Two practically relevant cases satisfy it. The first is monotonic spending: if spend always increases across the observation window, each observed spending level maps to a specific period, and therefore to a specific effectiveness. The paper observes that firms are reluctant to cut marketing budgets, so an approximately increasing trend is plausible. The second is a parent process, where spending and effectiveness are both determined by some third variable and the relationship between that variable and spending is invertible. Seasonality driving both effectiveness and budget timing is the obvious example.

Both conditions are made more likely by ordinary management practice. Autoregressive budgeting increases the chance of monotonicity. Flight paths that raise or cut spend by a fixed percentage each period induce monotonic series directly. Setting budgets according to known seasonal variation creates a parent process.

The functional forms underneath the argument

The 2024 paper spends its background section on the specific mathematical objects modern packages implement, and the detail is worth carrying because it explains why the identification problem is not an artefact of an unusual specification.

Nonlinear response has been hypothesised in the marketing literature since at least the 1950s, and modern practice has converged on two related functional forms. The more general is the Hill function, which the paper notes is equivalent to the log-logistic cumulative distribution function and to the ADBUDG function introduced in 1970. It carries an inflection point parameter and a shape parameter, and it is used in Meridian, Robyn and PyMC-Marketing alike. The paper records a documented weakness independent of anything else it argues: the Hill function is itself poorly identified, with different parameter combinations yielding effectively the same function over a finite range, a point established in the 2017 Google research paper. That is sometimes handled by fixing the shape parameter at one, producing what the 2017 paper calls the reach transformation, named for its use in modelling the relationship between reach and gross rating points in television.

Carryover has a parallel history. Distributed lag models including multiple lags of spending are rich but introduce many parameters per channel, raising both data requirements and overfitting risk. Stock variables collapse the lags into a weighted combination of past spending, with the weights following a known form. Geometric decay raises a single parameter to the power of the lag. Pascal decay uses the probability mass function of a negative binomial distribution, which permits a delayed peak in effectiveness. Under infinite lags with geometric weights, the model reduces to the Koyck specification, in which current outcome depends on current spend and on the previous period's outcome.

Time-varying effectiveness has its own literature stretching back to the 1970s, split between theory-based models deriving non-stationarity from a specific account of advertising response and unstructured models estimating time-dependent parameters directly. The simplest unstructured version is a linear specification with coefficients allowed to evolve, and the paper catalogues the smoothing devices used to make that estimable from one observation per period: random walks, state-space models incorporating independent variables in the evolution, and Bayesian dynamic factor models.

The paper also pauses on terminology, and the clarification is useful for reading any mix model documentation. The word dynamic has been applied in the literature both to carryover models and to time-varying coefficient models. The 2024 paper uses dynamic interchangeably with time-varying, and reserves the word carryover for lagged effects of any kind, whether captured through lagged variables, Koyck formulations or stock variables. The distinction matters because a model can carry both: revenue can be a nonlinear function of a stock variable, and a time-varying model can include carryover by letting the whole coefficient vector depend on time.

Why Gaussian processes

The empirical framework rests on Gaussian process priors, and the choice is instrumental rather than doctrinal. A Gaussian process is defined by a mean function and a covariance function, or kernel, such that any fixed set of inputs produces a multivariate normal distribution over the associated function values. The mean function encodes a prior expectation; the kernel governs how the function deviates from it, including smoothness and amplitude.

The paper uses a squared exponential kernel with two hyperparameters. The lengthscale, also called smoothness, determines how similar the outputs of nearby inputs are. The amplitude determines overall variability. Both are set to zero mean, which the paper describes as the most general assumption because it imposes no prior knowledge, while noting the framework could be generalised to use the Hill function as the mean in the nonlinear case.

Three reasons are given for the choice. Gaussian processes are a natural prior for unknown functions, as in the nonlinear model, and have been used successfully to capture parametric evolution in time-varying models. They are flexible, allowing nonparametric deviations from whatever the mean function encodes. And, critically for the argument being made, the kernel hyperparameters allow the authors to generate functions with controlled properties, which is what makes the simulation study possible at all.

The paper is explicit that the conflation question is orthogonal to this modelling choice. Splines are named as a natural alternative for unknown functions, with monotone variants noted as useful for mix modelling; generalised additive models are named as another; and state-space and time series models are named as the conventional tool for carryover and for coefficient evolution. Alternatives are demonstrated in an appendix. The identification problem does not belong to Gaussian processes.

For the modern multichannel data, the framework is extended with a time-varying intercept using a Trend-Season covariance function, formed as the sum of a squared exponential kernel and a periodic kernel capturing recurring patterns of a given cycle length. Summing kernels produces a valid kernel, which is what permits trend and seasonality to be captured in one parsimonious object.

How often it happens

The simulation study is the empirical core. Three data-generating processes were used: a nonlinear model built on a Gaussian process, a time-varying linear model built the same way, and a nonlinear model using specifically the Hill function that all three major open-source packages implement.

Six parameters were varied across three levels each, producing 2,187 settings, with 100 datasets generated per setting across 100 periods of observation. The varied parameters were the amplitude of the response function at 1, 2 and 5; its smoothness set as a ratio of the range of the independent variable at 0.1, 0.5 and 1; for the Hill process, the shape parameter at 0.5, 2 and 3.5 and the inflection point at 0.1, 0.33 and 1; the autocorrelation coefficient of spending at 0, 0.5 and 1; the transition variance of spending at 1, 5 and 10; noise in the response as a ratio of its deterministic component at 0.01, 0.1 and 0.2; and adstock carryover at 0, 0.3 and 0.8.

Ten observations were held out per simulation, and a setting was labelled conflated when the wrong model predicted the holdout at least as well as the true one. The paper describes this as an extremely high bar, since it requires the competing model to match or beat the true generating process rather than merely to fit acceptably.

The results are in two columns, without a stock variable and with high carryover.

True processAny conflation, no stockMajor conflation, no stockAny conflation, high stockMajor conflation, high stock
Nonlinear81%27%85%30%
Time-varying91%40%99%44%
Hill83%46%92%47%

Major conflation is defined as more than a quarter of a setting's simulations being conflated.

With adstock at 0.8, roughly 90 percent of all simulation settings exhibited non-zero conflation, and the time-varying process reached 99 percent. Applying a stock transformation to the spend variable - which every major mix modelling package does by default - measurably worsens the problem, because the transformation smooths the input series and smooth inputs are exactly what allows a time-varying coefficient to impersonate a curve.

A regression of conflation rate on the simulation settings sharpens the picture. High smoothness in the response function carried coefficients of 18.86 under the nonlinear process and 12.21 under the time-varying one. Noise in the response was the largest single driver, at 23.28 and 22.71 respectively for the high setting. High adstock contributed 1.59 and 2.49. The autocorrelation coefficient of spending mattered strongly under the time-varying process, at 13.38, and not at all under the nonlinear one, an asymmetry the authors read as confirming that the conditions producing conflation differ by process.

The noise result deserves stating plainly, because it is the one that transfers directly to practice: conflation is dramatically more likely when data are noisy, and the paper's own applications find that modern mix model data are very noisy.

What conflation costs

Two models predicting holdout data identically will not necessarily agree about what to do next, because optimising a budget is an interventional query rather than a predictive one. Holdout data is generated under status quo spending. Optimisation asks what happens under spending that has not been observed.

The mechanism is the same one that produced the conflation. A time-varying model approximates a nonlinear one by making a local linear approximation around observed spend, and learning a smooth coefficient requires reasonably smooth changes in spend. Optimisation intervenes on spend, breaking that smoothness, at which point the predicted coefficient is no longer a reasonable linear approximation to the true curve at the new spending level. Intervening in the other direction breaks whatever relationship existed between spend and effectiveness that allowed the nonlinear model to impersonate the time-varying one.

The simulated illustration is stark. On a dataset generated from a sigmoid response with autocorrelated spending and low noise, where both models fit equally well, the nonlinear model implied optimal expenditure of $12,072.60 while the time-varying model implied $7,451.50. The gap amounts to almost the entire range of the training data. The nonlinear model reasons that it is worth investing enough to get past the inflection point of the sigmoid; the log-log time-varying model has no inflection point and recommends a much lower spend given the implied elasticity at that moment.

Classic datasets, unresolved after fifty years

Two datasets long used in advertising response research were re-analysed. Both are monthly with two variables, spend and sales.

The dietary weight control product data from Bass and Clarke (1972) covers three years and reports sales in units. The Lydia Pinkham data from Palda (1965) covers 78 months of herbal medication sales from January 1954 to June 1960 in dollars, seasonally adjusted. For the Lydia Pinkham series, a stock transformation was applied with a lookback of 13 and a carryover of 0.7, the value derived by Bultez and Naert (1979) from their examination of the same data.

Four models were fitted to each: nonlinear via Gaussian process, time-varying via Gaussian process, nonlinear using the reach variant of the Hill function, and time-varying with a log-transformed stock variable.

DatasetNonlinearTime-varyingHillLog time-varying
Dietary weight control3.23 (2.22 to 5.58)2.97 (1.72 to 5.09)3.10 (2.59 to 4.44)2.55 (1.50 to 4.13)
Lydia Pinkham89.37 (79.41 to 104.88)83.21 (75.44 to 95.24)100.50 (91.68 to 114.97)88.92 (83.23 to 98.00)

Every posterior interval overlaps within each dataset. Standard model selection cannot separate them.

The allocation implications diverge in both cases. For the dietary weight control data, prices are not observed, so the optimisation was run across a grid of hypothetical prices. Across the majority of the price range, the time-varying model recommended higher spending than the nonlinear model, sometimes dramatically higher. For Lydia Pinkham, with the optimisation window placed far enough from the end of the series that carryover effects fully realise, the nonlinear model implied optimal spend of $90,127.63 and the time-varying model $78,333.87. The paper notes that the second figure is the lowest value the optimisation considered, since the observed post-test spending was lower than all prior periods, so the true time-varying optimum is likely lower still.

The nonlinear model on Lydia Pinkham exhibits an S-shaped response. The time-varying log-linear model implies a positive but decreasing elasticity across the same period. Those are different stories about the same 78 months, and the data do not choose between them.

Modern multichannel data

To test whether the problem persists in contemporary data, weekly sales and advertising figures were assembled from NielsenIQ Retail Scanner and Nielsen AdIntel covering 2018 and 2019, for four large national brands in separate categories: chocolate, pet food, coffee and beer.

Three features of modern data required modelling changes. Multiple advertising channels were already handled by the framework. Trend and seasonality were handled by adding a time-varying intercept with a covariance function combining a squared exponential kernel and a periodic kernel, plus dummies for major United States holidays. Promotions, unobserved except as spikes in spending, were captured with a constructed dummy variable flagging sparse burst spending on rarely used channels and unusual spikes on the main channels. Adstock was applied with decay set to 0.3, described as consistent with industry practice, and 13 lags, and the adstocked data was logged.

BrandNonlinear RMSETime-varying RMSE
Chocolate368,766 (207,634 to 601,163)387,623 (207,350 to 671,246)
Pet food48,438 (27,920 to 83,249)41,825 (24,631 to 65,232)
Coffee722,076 (509,527 to 919,820)733,269 (490,785 to 994,303)
Beer443,817 (236,730 to 911,732)423,085 (207,914 to 802,740)

The point estimates split two-two, with nonlinear ahead on chocolate and coffee and time-varying ahead on pet food and beer. Every interval overlaps. In all four cases, standard model selection practice would be unable to identify the correct specification.

The paper also records an observation that will be familiar to anyone who has read a mix model output: across all four modern datasets, a large share of the variation in revenue is captured by the trend and seasonality component rather than by marketing spend at all.

The coffee brand, and a number attached to being wrong

The coffee case study puts a currency figure on the consequence. The exercise redistributed the actual advertising budget for the last week of September 2019 across three channels, varying allocations in 5 percent increments, discarding any allocation placing a channel outside its historical spending range, computing the implied adstocks, and predicting sales for the target week and the following 13 weeks to capture carryover. Total spending was held constant by construction.

ModelDigitalNetwork TVSpot TVCost of conflation
Nonlinear15%50%35%$227,000
Time-varying15%85%0%$61,000

Both models agree on digital at 15 percent, which was the maximum observed in the data and therefore a binding constraint rather than a finding. They disagree completely on television. The nonlinear model allocates 35 percent to spot television. The time-varying model allocates nothing.

The reasoning behind each is visible in the estimated components. The nonlinear model finds a peak in the effect of network television followed by a declining effect, so it recommends spending some but not too much there and diverts the remainder to spot. The time-varying model finds network television's elasticity higher than spot television's in almost every period, and allocates accordingly.

The final column prices the error. If the nonlinear model is correct and the time-varying allocation is followed, the loss over the following 14 weeks is $227,000. If the time-varying model is correct and the nonlinear allocation is followed, the loss is $61,000. The paper emphasises the narrowness of the exercise: only channel allocation was varied, not total spend, and only in a single period. The economic consequence of misspecification is therefore a lower bound rather than an estimate of annual exposure.

Breaking the tie deliberately

The paper's constructive contribution is a pair of spending policies designed to make the two model families disagree, so that ordinary holdout metrics can then identify the correct one.

The first is a maximal separation test. Given both models estimated on historical data, a range of candidate spending levels is considered, both models predict the next period's outcome at each candidate, and spend is set to the level that maximises the absolute difference between the two predictions. After the period runs, whichever prediction was closer to the realised outcome favours that model.

Applied to two simulated examples with 48 periods of history, the posterior distributions of the two models' in-sample errors overlapped completely before the test began, and separated after two periods of testing, with the true model always carrying the lower error. Separation happened marginally faster when the true process was time-varying, consistent with the theory, since the conditions under which a nonlinear model can approximate a time-varying one are stricter.

The dynamics of the test are informative in themselves. When the true process is nonlinear, the maximally separating pattern alternates between high and low spending, because the time-varying model's regularisation implies a locally linear response and a linear approximation cannot track an oscillating response to two extreme spending levels. When the true process is time-varying, the maximally separating level stays stable for some time, exploiting the fact that the same spending level produces different responses as the elasticity evolves - something a static nonlinear model cannot reproduce. The procedure is adversarial: as the wrong model adapts to the new spending pattern, the test finds a new value to probe.

The second policy is a simplification of the first, and the one more likely to reach a media plan. The seesaw test alternates spending between a high and a low level, one period at a time. It resembles what practitioners call a bump-up test, and it also resembles pulsing, an advertising strategy shown in earlier literature to be both common and sometimes optimal. Applied to the same two simulated processes, the errors of the two models diverged after a few intervention periods.

The authors note a limitation. Under strong carryover with long lags, the smoothing induced by the stock variable may limit how much variability a seesaw pattern can create in a short window, in which case direct implementation of the maximal separation test is the more effective option.

Where the two papers meet

Both papers arrive at experiments, from opposite directions.

Heusch argues that experiments are necessary because observational estimation cannot achieve causal identification, and that current practice wastes most of the experimental information it collects. Dew, Padilla and Shchetkina argue that experiments are necessary because observational data cannot distinguish two model families that fit identically and prescribe differently, and they design specific spending interventions to force the distinction.

The two requirements are not the same and are not in tension. One asks for tests long enough and varied enough in spending level to trace a response curve and observe decay. The other asks for spending patterns that deliberately break the local smoothness that lets one model impersonate another. A seesaw test alternating between high and low levels within a short window and a four-week go-dark test at a fixed level are different designs. A programme that satisfies both would be more demanding than either.

There is also a substantive gap between them worth naming. Heusch assumes parameters are constant across experimental periods, and cites Dew et al. as the reference for the alternative. Dew, Padilla and Shchetkina show that time-varying effectiveness is not merely possible but frequently indistinguishable from the nonlinearity that mix models assume. If effectiveness genuinely varies over time in a given account, the structural estimates recovered from pooled geo-tests represent some form of average across the experimental windows, which the Heusch paper acknowledges. What neither paper resolves is how a practitioner would know which situation applies before committing to a design.

The earlier paper also leaves a thread the later one picks up. Dew, Padilla and Shchetkina explicitly state that they have not explored how endogeneity intersects with the conflation issue, noting that the problems they document exist even with purely exogenous spending and would likely be compounded by endogeneity and its corrections. The Heusch paper documents an endogeneity failure of a factor of 2.5 that survives oracle controls. The interaction of the two remains unexamined.

Both papers converge on one practical observation about current tooling. Both note that Meta's Robyn and Google's Meridian already direct users toward incrementality tests, and both argue the tests are being used for less than they could be. The 2024 paper puts it as an additional reason experiments matter: a carefully designed experiment can disentangle nonlinear from time-varying effects and identify the right specification, not only mitigate endogeneity.

Why this matters to media buyers

The commercial stakes attach to the allocation decision rather than to the reported number.

A mix model exists to answer where the next unit of budget goes. Both papers demonstrate that models fitting observed data acceptably can be wrong about that answer in specific, quantifiable ways. The 2024 paper prices one instance at $227,000 across 14 weeks for a single channel reallocation at a single brand. The 2026 paper documents a channel return misreported by a factor of 2.5, in the direction that attracts budget.

That lands on a discipline already under scrutiny. IAB's State of Data 2026 report, published February 7, 2026, found up to 75 percent of buy-side decision-makers rating attribution, incrementality tests and marketing mix models as underperforming on rigour, timeliness, trust and efficiency. Incubeta research released on May 6, 2026 recorded 70.4 percent of leaders confident their budgets were deployed effectively while 41.6 percent conceded waste. Survey work covered in November 2025 found 46.9 percent of marketers planning to increase marketing mix modelling investment over the following year, the highest priority among measurement methods surveyed, with 52 percent already using incrementality testing and experiments.

Investment is rising into a method whose identification properties two independent research groups have now questioned in print. That is not an argument against the method. It is an argument about what the outputs can carry.

The governance dimension compounds it. A PPC Land analysis published April 2, 2026 examined the structural dynamic when a platform builds the tool guiding budget allocation, using Meta's Robyn as the case. CIMM returned to the same territory on July 31, 2026, warning that default priors and embedded assumptions in widely adopted open-source frameworks can reconfigure data asymmetry rather than remove it. The Heusch result adds a technical point to that argument: the seasonal controls the frameworks ship as defaults were the exact specification that produced the 2.5-fold error, and adjusting the priors made it worse rather than better.

Both frameworks are open source, which is what makes this kind of scrutiny possible at all. A closed vendor model producing 10.61x would not expose the specification that produced it. The papers' criticisms are directed at published methods precisely because those methods are published.

Data supply is moving the same way. Amazon's Marketing Mix Modeling API reached general availability in May 2026, giving advertisers programmatic access to aggregated advertising and retail signals across 14 countries. IAB published a white paper on April 7, 2026 arguing that standard mix modelling is structurally misaligned with retail media measurement and causing brands to undervalue the channel. More data flowing into a model with an identification problem does not fix the identification problem, which is the specific point the oracle specification was constructed to make.

The vendor landscape has been consolidating around the same claims. Prescient AI announced in July 2025 what it described as the first mix model built from scratch since the 1960s, positioning against open-source foundations it characterised as resting on outdated mathematical assumptions. IAB Australia published a vendor landscape in September 2025 documenting the range of commercial approaches. Neither paper considered here evaluates any commercial product, and neither claims that all implementations share the documented failure. What they establish is that the failure is a property of the estimation problem rather than of any one codebase, which means the burden of demonstrating that a given implementation escapes it sits with whoever makes that claim.

Meridian's own trajectory suggests the platform is aware of the gap. The framework moved from a standalone Python package to an enterprise stack integrated into Google Analytics 360 on May 20, 2026, with GeoX added as a geographic experimentation layer feeding results into the model. The Heusch paper's argument is not that such a layer is unnecessary. It is that feeding a single aggregate figure from that layer into the model wastes the temporal structure the experiment generated.

The line that will get quoted

The 2024 paper closes with a recommendation addressed to managers rather than to econometricians, and states that managers "should not blindly trust existing MMM solutions", adding that care is required in interpreting results because the aggregated data used for estimation often cannot support identification of complex effects.

The 2026 paper's equivalent is quieter and arguably harder. Its practical implication is that many firms already run geo-experiments and discard most of the information those experiments generate, and that calibrating a complete functional form is available from the same experimental investment already being made.

Whether either claim survives contact with real advertiser data is the open question in both cases. One paper runs on synthetic data with known parameters. The other runs on real data where the parameters are unknown, which is why it can demonstrate that two models fit equally well but cannot demonstrate which one is right. The gap between those two positions is the reason experiments keep appearing at the end of both arguments.

Timeline

Summary

Who: Niklas Heusch, whose contact address at the head of the paper carries a zalando.de domain, authored the August 2026 paper. Ryan Dew and Anya Shchetkina of the Wharton School and Nicolas Padilla of London Business School authored the 2024 paper it builds on. The findings concern advertisers, agency measurement teams and data scientists running marketing mix models, and reference the open-source frameworks maintained by Google, Meta and the pymc-marketing project.

What: The 2026 paper proposes estimating the complete set of mix model parameters - adstock decay, saturation curvature and effectiveness - directly from geo-experiment time series by differencing treatment and control outcomes, and demonstrates on synthetic data that a conventionally specified mix model reported 10.61x return on ad spend against a true 4.20x, that the same model handed the true confounders still reported 8.41x, and that the structural approach recovered 4.31x from two experiments and 4.14x from four. The 2024 paper demonstrates that nonlinear and time-varying effects are frequently not separately identifiable from standard mix model data, with conflation rates reaching 99 percent of simulation settings under high carryover, and prices one instance of the resulting misallocation at $227,000 across 14 weeks.

When: The structural estimation paper was posted to arXiv on 21 August 2026 as 2608.21128v1. The identification paper was posted on 14 August 2024 as 2408.07678v1 with a title-page date of August 15, 2024.

Where: Both papers are preprints on arXiv, the first filed under Applications in statistics and the second under econometrics. The 2026 results rest entirely on synthetic data simulating an online retailer across 156 weeks with three channels. The 2024 results combine a 2,187-setting simulation study with re-analysis of the Lydia Pinkham and dietary weight control datasets and four brands built from NielsenIQ Retail Scanner and Nielsen AdIntel data for 2018 and 2019.

Why: Marketing mix modelling has become the primary cross-channel measurement method as user-level attribution degraded under privacy regulation, and investment in it is rising. Both papers identify failures that are properties of the estimation problem rather than of any single implementation, and both conclude that experimental variation - designed for structural estimation rather than reduced to a single lift figure - is what the models need to produce allocation decisions that hold under intervention.