Endogeneity is the condition in which an explanatory variable in a statistical model is correlated with the model's error term, meaning the variable carries information the model was supposed to hold constant. In advertising, the explanatory variable is usually spend or exposure, and the error term contains everything the analyst could not measure: latent demand, brand momentum, a consumer's pre-existing intent to buy. When the two move together, the estimated effect of advertising absorbs the effect of the thing that caused the advertising. The coefficient stops describing causation and starts describing coincidence.

The term names the largest source of measurement error in the industry. It explains why a platform dashboard, a multi-touch attribution model and a marketing mix model can all be internally consistent, well fitted and wrong in the same direction. It also explains why randomised experiments sit at the top of measurement hierarchies: randomisation is not a superior statistical technique but the only reliable way to break the correlation.

The mechanics of a biased coefficient

The canonical model regresses an outcome on advertising: sales equals a constant, plus a coefficient multiplied by spend, plus an error term. Ordinary least squares recovers the true coefficient only if the error term is uncorrelated with spend, a condition econometricians call exogeneity. Break that condition and the estimate absorbs the covariance between the two, producing a bias whose sign follows the sign of the correlation.

In advertising the correlation is almost always positive, so the bias is almost always upward. Four mechanisms produce it, and they operate simultaneously.

Targeting selection. Ad delivery systems are built to find users with high predicted conversion probability. Exposure is therefore assigned to people who were already more likely to buy, which is precisely the variable the model cannot observe. The exposed group and the unexposed group differ before a single impression is served.

Activity bias. Users who are online at a given moment are simultaneously more likely to see an ad and more likely to do anything else measurable, including searching for a brand. Randall Lewis, Justin Rao and David Reiley coined the phrase at the World Wide Web conference in Hyderabad in 2011, showing that a control group given a placebo exposure produced almost the same spike in brand-relevant search as the treated group.

Simultaneity. Budgets respond to expected demand. Spend rises before Black Friday because sales are expected to rise, not the other way round, and quarterly budgets set as a share of the previous quarter's revenue embed that quarter's shocks. Causation runs in both directions between the two variables the model treats as one dependent and one independent.

Algorithmic feedback. Auction-time bidding reads more than twenty signals and adjusts continuously to realised performance, which means spend chases the random component of weekly outcomes. Nothing an analyst can hold constant absorbs it, because the thing being tracked is noise.

The concept predates advertising by roughly a century. Its formal treatment begins with the identification problem in economics: observed prices and quantities trace the intersection of supply and demand, not either curve. Philip G. Wright's 1928 study of oil tariffs contains, in Appendix B, the first published instrumental variable estimator, a method for isolating variation in one equation using a variable that shifts it without touching the other. Olav Reiersol supplied the name "instrumental variables" in 1941. Trygve Haavelmo's 1943 paper on simultaneous equations and his 1944 probability monograph established that interdependence biases least squares, work carried forward by the Cowles Commission under Jacob Marschak and Tjalling Koopmans. Henri Theil developed two-stage least squares in 1953.

Digital advertising rediscovered the problem in the 2010s, with data rather than theory. The eBay experiments reported by Thomas Blake, Chris Nosko and Steven Tadelis in Econometrica in 2015 remain the reference case. eBay suspended brand keyword bidding on Yahoo and MSN in March 2012 while continuing to pay on Google as a control, and found that 99.5 percent of forgone paid click traffic was recaptured by organic search. A second experiment cut non-brand bidding across 68 of 210 designated market areas for 60 days. Ordinary least squares run on the pre-experiment data implied a return on investment of 4,173 percent without geographic and date controls and 1,632 percent with them. The experimental estimate was minus 63 percent, with a 95 percent confidence interval running from minus 124 to minus 3.

Two orders of magnitude separated the observational reading from the causal one, in a dataset with no tracking loss.

What richer data does not fix

The obvious rejoinder is that eBay's regression was crude. Later work removed that defence.

Brett Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky compared observational estimates against 15 randomised experiments at Facebook covering 500 million user-experiment observations and 1.6 billion impressions, publishing in Marketing Science in March 2019. Matching and regression methods failed to reproduce the experimental effects even after conditioning on extensive demographic and behavioural variables. In half the studies, the estimated percentage increase in purchase outcomes was off by a factor of three.

Gordon, Robert Moakler and Zettelmeyer extended the test to 663 Facebook experiments described by more than 5,000 user-level features, published in Marketing Science in July 2023. Double machine learning outperformed stratified propensity score matching, and neither worked. Median absolute error in measured lift reached 115 percentage points for upper-funnel outcomes, 107 for mid-funnel and 62 for lower-funnel, against median experimental lifts of 28, 19 and 6 percent respectively. The error exceeded the effect.

A preprint posted to arXiv on 21 August 2026 by Niklas Heusch made the same argument on the aggregate side. A marketing mix model using the seasonal controls that Meridian, Robyn and pymc-marketing ship as defaults reported 10.61x return on ad spend for a simulated paid search channel whose true return was 4.20x. Handed the true confounders as a diagnostic, it still read 8.41x. Looser priors moved the figure to 11.38x rather than down.

Why it matters for media buyers

Bias in a stated direction is a budget decision, not a reporting quirk. A channel reported at twice its true return attracts marginal spend priced against a response curve that does not exist, and the money comes out of channels whose measurement happens to be less flattering.

IAB and IAB Europe published commerce media incrementality guidelines in November 2025 requiring credible counterfactuals, and naming selection bias in observational data as a critical source of model bias. IAB's State of Data 2026 report, published on 7 February 2026, found up to 75 percent of buy-side decision-makers rating attribution, incrementality tests and mix models as underperforming. LiveRamp simulations from July 2026 showed identity precision errors reading a 25 percent true lift as 6.8 percent, a reminder that measurement error and endogeneity push in opposite directions and can mask each other.

Access to the correction has widened. Google cut the minimum incrementality experiment budget from thresholds approaching $100,000 to $5,000 on 11 November 2025, reporting that most advertisers run roughly one to two studies a year. Geo-split designs of the kind TikTok documents as Geo Lift Tests follow the framework Jon Vaver and Jim Koehler published at Google in 2011.

Limitations and disputes

Randomisation removes the bias but not the difficulty. Lewis and Rao showed in the Quarterly Journal of Economics in November 2015, using 25 field experiments with major United States retailers and brokerages representing $2.8 million of spend, that the median confidence interval on return on investment ran more than 100 percentage points wide. Individual sales carry a coefficient of variation around 10, and informative experiments can require more than ten million person-weeks. An unbiased estimate that cannot exclude zero settles nothing.

The direction of bias is not universal. Gordon and co-authors found observational methods generally overstating effects but significantly understating them in some cases, which means a practitioner cannot apply a haircut and move on.

Experiments carry their own threats. Geographic spillover violates the assumption that control-area advertising does not reach treated consumers. Divergent delivery, where an algorithm serves treatment and control cells to different people, contaminates platform-run split tests. And the party running the experiment is frequently the party selling the media, a conflict PPC Land examined on 2 April 2026 in the case of Meta's Robyn. CIMM warned on 31 July 2026 that default priors in open-source frameworks can reconfigure data asymmetry.

There is also a live methodological dispute about whether experiments should be summarised or modelled. Existing calibration reduces a test to a single lift figure fed into a mix model as a prior. The Heusch paper argues that differencing treatment and control time series recovers adstock, saturation and effectiveness directly, and that collapsing the series discards the dynamics. That approach requires tests spanning different spending levels, which costs more than most advertisers currently spend on measurement.

What it is not

Omitted variable bias is one route to endogeneity rather than a synonym. A missing confounder correlated with both spend and sales produces it, but so do simultaneity and measurement error in the regressor, neither of which involves a missing variable.

Selection bias describes non-random assignment of exposure. It is the dominant form endogeneity takes in user-level measurement, while aggregate models more often inherit it from budget setting.

Multicollinearity occurs when explanatory variables move together, inflating standard errors. It makes estimates imprecise without making them biased. Endogeneity does the opposite: it can produce a tight interval around the wrong number.

Endogenous variables in a model specification are simply those determined inside the system, as distinct from exogenous controls determined outside it. The label is descriptive. Endogeneity is the failure that follows when a variable treated as exogenous is not.

Timeline

  • 1928 - Philip G. Wright publishes The Tariff on Animal and Vegetable Oils, whose Appendix B contains the first published instrumental variable estimator
  • 1941 - Olav Reiersol introduces the term "instrumental variables"
  • 1943 - Trygve Haavelmo publishes on the statistical implications of simultaneous equations, followed by his 1944 probability monograph
  • 1953 - Henri Theil develops two-stage least squares
  • March 2011 - Lewis, Rao and Reiley present activity bias at the World Wide Web conference in Hyderabad
  • 2011 - Vaver and Koehler publish the geo-experiment framework for measuring ad effectiveness
  • March 2012 - eBay suspends brand keyword advertising on Yahoo and MSN, retaining Google as a control
  • January 2015 - Blake, Nosko and Tadelis publish the eBay experiments in Econometrica
  • July 2015 - Lewis and Rao's statistical critique of advertising measurement reaches advance access at the Quarterly Journal of Economics
  • March 2019 - Gordon, Zettelmeyer, Bhargava and Chapsky publish the 15-experiment Facebook comparison in Marketing Science
  • July 2023 - Gordon, Moakler and Zettelmeyer publish the 663-experiment study of double machine learning and propensity score matching
  • 14 August 2024 - Dew, Padilla and Shchetkina post "Your MMM is Broken" to arXiv
  • 3 November 2025 - IAB and IAB Europe release commerce media incrementality guidelines
  • 11 November 2025 - Google lowers its minimum incrementality experiment budget to $5,000
  • 7 February 2026 - IAB State of Data 2026 records buy-side dissatisfaction with existing measurement methods
  • 21 August 2026 - Heusch posts the structural estimation paper proposing recovery of mix model parameters from geo-experiment time series

Summary

Who: Econometricians including Philip G. Wright, Trygve Haavelmo and Henri Theil established the concept and its remedies; Randall Lewis, Justin Rao, David Reiley, Thomas Blake, Chris Nosko, Steven Tadelis, Brett Gordon, Florian Zettelmeyer, Robert Moakler and Niklas Heusch documented it in digital advertising. It affects advertisers, agency measurement teams, platform data scientists and the vendors selling attribution and mix modelling.

What: Correlation between advertising spend or exposure and the unobserved determinants of the outcome being measured, which biases estimated advertising effects, usually upward. It arises from targeting selection, activity bias, budgets responding to expected demand, and automated bidding responding to realised performance.

When: Formalised in econometrics between 1928 and 1953. Documented empirically in digital advertising from 2011 onward, with the eBay results published in 2015, the Facebook comparisons in 2019 and 2023, and an aggregate-model demonstration posted in August 2026.

Where: Present in platform-reported returns, multi-touch attribution, and marketing mix models alike. It is broken only where exposure or spend is assigned at random, whether by user-level holdout, placebo and ghost-ad designs, or geographic split tests.

Why: Because budget allocation follows measured return. A channel whose reported return is inflated by the demand that caused its spend attracts money that would earn more elsewhere, and no volume of additional observational data corrects the error.