Statistical modelling in advertising is the production of a number that was never directly observed, by fitting a mathematical relationship to data that was. A panel of 42,000 homes stands in for the whole television population. A purchase that no cookie recorded is inferred from the rate at which similar, cookie-visible purchases occurred. A revenue figure is split into the share each media channel contributed, using only weekly spend and sales totals. In every case the reported figure is an estimate carrying an error range, not a count.

It exists because measurement coverage has never matched measurement demand, and the gap has widened. Census counting, meaning a log entry for every impression, was digital advertising's founding advantage over television, which had relied on samples since audience research began. Browser restrictions on third-party cookies, mobile identifier permissions and consent requirements have removed enough of that census that digital measurement now faces the sampling problem broadcast measurement solved decades earlier, and has borrowed the same statistical apparatus to answer it.

How a modelled number is built

The most widely deployed example is conversion modelling, and Google's advertiser documentation sets out the arithmetic in four steps. Ad interactions are first separated into a group with an observable link to a conversion and a group without one. The observed group is then divided into subgroups sharing characteristics such as location, time of day and browser, with conversion rates calculated for each; the documentation's own example contrasts conversions observed in the morning in France with the evening rate in the same market. Unobserved interactions are sorted into those same subgroups. The known rates are then applied to produce estimated links, which are added to observed conversions in reporting and passed into automated bidding.

Two properties of that process matter more than the mathematics. The subgroups are the model's assumption: it holds that unobserved users behave like observed users sharing those attributes. And the accuracy test is internal. Google describes holdback validation, in which a slice of observed conversions is deliberately withheld, modelled, and the estimate compared against the known answer.

Projection works differently. Panel and sample measurement takes activity from a recruited group and scales it to a population using a universe estimate, a stated count of the people or households the sample represents. The Media Rating Council's Digital Audience-Based Measurement Standards, published in final form in December 2017, require that universe figures come from independent industry or governmental sources, that weighting methods be supported by empirical study, and that standard errors around sample-based projections be disclosed. The standards set a general materiality threshold of 5 percent of reported activity by reporting break. Where a measurement service enriches its records by joining them to an outside dataset, the fields used as links must have demonstrable power, meaning a statistically established ability to explain differences in media consumption.

Nielsen's national television currency shows the machinery running at scale. Big Data plus Panel merges a 42,000-home panel of more than 100,000 people with device inputs from roughly 45 million households and 75 million devices. On 31 August 2026 Nielsen altered seven metrics inside that product, among them a Household Demographic Assignment Model, integrated weighting, an automatic content recognition tuning adjustment and a co-viewing enhancement. Each is a modelling decision, and each moves reported audiences without any change in viewing.

Four jobs modelling does

Gap filling estimates values the measurement system could not capture. Google introduced behavioural modelling in Google Analytics in 2022, estimating the conduct of users who declined analytics cookies from those who accepted. The eligibility thresholds are published: consent mode must run on every page, the property must record at least 1,000 daily events carrying a denied analytics storage flag for seven days, and at least 1,000 daily users must be sending consented events. Below those volumes the model does not run, which makes gap filling a function of site scale. Amazon DSP applies the same logic to conversions on anonymous inventory.

Projection scales a sample to a population. It is how television and cross-media currency is produced, and how Google generates store visit figures by extrapolating from a subset of signed-in users with location history enabled.

Decomposition splits an outcome among causes. Marketing mix modelling is the dominant form: Meridian, Google's open-source Bayesian framework, opened globally on 29 January 2025 and added non-media variables such as pricing and promotions that September. Data-driven attribution performs a narrower version at touchpoint level.

Classification assigns a label. Invalid traffic detection sorts requests into valid and invalid; the MRC's updated detection and filtration standards, documented in April 2022, formalised requirements for machine learning and for sampling approaches. Probabilistic device graphs sit here too, inferring that several devices belong to one person from signals including type, operating system and network.

Origin and evolution

Sampling arrived in media measurement long before digital advertising, and person-level television measurement dates to the 1987 people meter. Digital inherited a census and lost it gradually.

Google published the research underpinning its mix model in 2017, covering Bayesian methods with carryover and shape effects and geo-level hierarchical modelling. Modelled conversions appeared in Google Analytics the same year, reached Google Ads in July 2019, and arrived in Display and Video 360 in 2020, where the company told advertisers to expect a low double-digit increase in reported conversions with no detail on how the model worked.

Regulation accelerated the shift. Consent Mode gave Google a mechanism for modelling conversions lost to declined consent, and version 2 became a condition of measurement in Europe when Google began disabling personalisation, remarketing and conversion tracking for non-compliant advertisers on 21 July 2025. Aggregate methods returned in parallel. Nielsen ended stand-alone panel ratings in January 2025 and launched Big Data plus Panel for the September 2025 season. Google cut the minimum budget for an incrementality experiment to $5,000 on 11 November 2025, citing a switch to Bayesian methodology that reached conclusions on less data.

Why it matters for marketers

Modelled figures no longer sit in a footnote. They enter the conversion column, feed automated bidding and set the budget split. When the estimate degrades, spending degrades with it. Advertisers whose consent banners collected choices without transmitting them to Google's tags recorded conversion collapses of around 90 percent, with roughly 40 percent of attribution data recoverable and the rest permanently lost, because the model had nothing observed to extrapolate from.

Reported confidence is low. IAB's State of Data 2026, published on 7 February 2026, found up to 75 percent of buy-side decision-makers rating attribution, incrementality tests and mix models as underperforming on rigour, timeliness, trust and efficiency.

Limitations and disputes

The structural criticism is that a modelled number cannot be independently verified by the party relying on it. Advertisers do not know what share of a reported total is modelled, a complaint on record since 2020 and unresolved.

Specification errors carry prices. A 2026 preprint by a Zalando researcher found mix models overstating paid search return on ad spend by a factor of 2.5, in the direction that attracts budget, with an earlier study pricing one misallocation at $227,000 across 14 weeks at a single brand. The Coalition for Innovative Media Measurement warned on 31 July 2026 that default priors embedded in widely adopted open-source frameworks can reconfigure data asymmetry rather than remove it.

Method disagreement is visible in outputs. Co-viewing estimates for the same household watching the same programme have differed between providers, counted as 1.5 viewers by one and 1.2 by another, because the arithmetic differs rather than the behaviour. The Video Advertising Bureau warned video buyers in 2026 that automatic content recognition data requires opt-in and set-top box data skews toward wealthier homes, biases that survive projection.

Standards bodies have drawn one hard line. The MRC standards forbid a census measurement organisation from reporting unique users purely through algorithms or modelling not at least partially traceable to information obtained directly from people. Modelling may extend observation; it may not replace it entirely.

What it is not

Machine learning names a family of fitting techniques, not the task. A mix model estimated by Bayesian regression is statistical modelling; so is a neural network classifier. The distinction is method, not purpose.

Attribution allocates credit among touchpoints. Last-click attribution is a rule, involving no estimation at all, while data-driven attribution is a model. The two are frequently conflated.

Incrementality testing randomises exposure to observe a counterfactual rather than estimate one. Experiments are used to calibrate models, notably through geographic tests such as Meridian GeoX, announced on 5 May 2026, but the two are different epistemic acts.

Deterministic matching joins records on a shared identifier. Probabilistic matching estimates the same link. Only the second is modelling.

Recent developments

Modelling infrastructure is being packaged rather than published. Google placed Meridian inside Google Analytics 360 at Google Marketing Live in May 2026 alongside a Gemini-powered predictive metric, and Amazon moved its marketing mix modelling API to general availability across 14 countries the same month. Both moves shorten the distance between a platform's data and the model allocating budget to it.

Timeline

  • 1987: people meters introduce person-level television measurement
  • 2017: Google Research publishes Bayesian and geo-level hierarchical mix modelling papers; modelled conversions appear in Google Analytics
  • December 2017: MRC publishes final Digital Audience-Based Measurement Standards version 1.0
  • July 2019: modelled conversions reach Google Ads
  • 2020: modelled conversions reach Display and Video 360
  • April 2022: MRC updates Invalid Traffic Detection and Filtration Standards, formalising machine learning requirements
  • 2022: behavioural modelling launches in Google Analytics, reaching realtime reports in December of that year
  • March 2024: Google unveils Meridian on limited availability
  • 24 January 2025: Nielsen ends stand-alone panel-based television ratings
  • 29 January 2025: Meridian opens globally
  • 21 July 2025: Google begins disabling conversion tracking for non-compliant EU advertisers
  • 2 September 2025: Nielsen launches Big Data plus Panel for the 2025 season
  • 30 September 2025: Meridian adds non-media variables and channel-level contribution priors
  • 11 November 2025: Google cuts minimum incrementality experiment budget to $5,000
  • 7 February 2026: IAB publishes State of Data 2026
  • 5 May 2026: Google announces Meridian GeoX and Meridian Studio
  • May 2026: Amazon moves its marketing mix modelling API to general availability
  • 31 July 2026: CIMM warns on default priors in open-source frameworks
  • 31 August 2026: Nielsen activates seven methodology changes in its national currency

Summary

Who: Measurement providers including Nielsen, Comscore and the verification vendors; platforms including Google, Amazon, Meta and TikTok that model conversions inside their own reporting; standards bodies including the Media Rating Council, the IAB and CIMM; and the advertisers, agencies and publishers transacting against the resulting figures.

What: The estimation of advertising quantities that were not directly observed, from data that was, spanning conversion and behavioural modelling, panel projection to population universes, marketing mix and attribution decomposition, and probabilistic classification of traffic, devices and demographics.

When: Sampling and projection long predate digital advertising, with person-level television measurement dating to 1987. Digital adoption accelerated from 2017, when Google published its mix modelling research and introduced modelled conversions, through the MRC's December 2017 digital audience standards, Consent Mode enforcement in July 2025, Nielsen's Big Data plus Panel launch in September 2025 and the packaging of mix models into platform products through 2026.

Where: Across search, social, display, retail media, connected television and linear television, in platform reporting interfaces, accredited currency products and open-source modelling frameworks.

Why: Measurement coverage has fallen below what buying and selling require, as cookies, device identifiers and consent were withdrawn. Modelling restores completeness at the cost of verifiability. The trade-off is contested: platforms report validated accuracy while trade bodies record up to 75 percent of buy-side decision-makers rating modelled measurement as underperforming, and published research has priced specific specification errors at multiples of the true figure.