A holdout study is a randomised experiment in which part of an advertiser's addressable audience is deliberately prevented from seeing a campaign, so that the outcomes of the excluded group stand in for what would have happened without the advertising. The excluded group is the holdout, or control. The exposed group is the treatment. The difference between the two, adjusted for group size, is the campaign's measured effect.
The design exists because no volume of tracking answers the question it addresses. Conversion tracking records that a purchase followed an impression; it cannot establish that the impression caused it. Retargeting and branded search are the standard illustrations, since both concentrate spend on people already moving toward a transaction. A holdout study replaces that inference with a comparison against a group drawn at random from the same population.
How the design is assembled
Three decisions define a holdout study: the unit of randomisation, the size of the holdout and the treatment applied to the control.
The unit is usually the logged-in user, the device, the household or the geographic region. User-level assignment is cleanest statistically and most demanding operationally, since it requires the platform to recognise the same person across sessions. Meta assigns a random 64-bit value to each test, concatenates it with a user identifier, applies a Murmur hash and maps the output into buckets allocated to conditions in the proportions the advertiser specified. Regional designs sidestep identity by splitting markets instead of people, an approach Google formalised in a 2011 paper by Jon Vaver and Jim Koehler that remains the reference for geographic experiments.
Holdout size is a direct trade. Google permits holdbacks between 1% and 50% of eligible traffic for user-based Conversion Lift, and its documentation is explicit that larger holdouts collect more control data while forgoing more conversions. Smaller holdouts cost less but need longer flights. Practitioners commonly run control groups at 5% to 10% of test size, which is also why many studies end up underpowered.
The third decision is the least visible and the most consequential. Early designs served the control group a public service announcement in place of the advertiser's creative, which kept exposure symmetrical but meant paying for impressions carrying no commercial message. A 10% holdout under that model consumes roughly a tenth of the budget on non-campaign delivery. Platforms now suppress instead. In a Meta Lift test, control users pass through the identical delivery pipeline; the campaign's ads are removed immediately before display and the next-best ad in the auction is shown. Because both groups traverse the same ranking and pacing logic, algorithmic allocation affects them symmetrically. Google's geographic implementation clusters geo-targets into units built to limit how often someone sees an ad in one region and converts in another.
What the study returns
Absolute lift, reported by Google as incremental conversions, is treatment conversions minus control conversions. Relative lift divides that figure by control conversions, so it can theoretically run from minus one to infinity and inflates sharply when baseline activity is low, a distortion Google's own documentation flags as a source of misreading across studies. Incremental cost per action divides spend by incremental conversions; incremental return on ad spend divides incremental conversion value by spend.
Each of those numbers is an estimate with an interval around it, and both major platforms now report Bayesian intervals rather than frequentist significance. Meta samples the posterior distribution of per-user incremental conversions for the narrowest span containing 90% of its area, and labels a result highly conclusive when more than 90% of the posterior sits above zero. Google reports study power, its estimate of the certainty a study will reach a conclusive answer, as a range from 50% to 95% in five-point increments, with budget guidance when a configuration falls below 90%.
Duration follows conversion lag rather than convention. Google allows studies as short as seven days, recommends at least fourteen, and reports up to a 17% drop in absolute lift for long-lag businesses running shorter flights. Most advertisers, by its own account, run one to two studies a year.
Origin and evolution
Randomised advertising experiments predate digital media, but their economics were examined seriously only once platforms could run them at scale. Randall Lewis and Justin Rao published the decisive assessment in the Quarterly Journal of Economics in November 2015, analysing twenty-five field experiments with United States retailers and brokerages representing $2.8 million of ad spend. The median confidence interval on return on investment exceeded 100 percentage points, and individual sales were volatile enough that informative experiments could require more than ten million person-weeks.
The response was to make experiments cheaper rather than larger. Garrett Johnson, Randall Lewis and Elmar Nubbemeyer published the ghost ads method in the Journal of Marketing Research in December 2017, having circulated it from 2015. Rather than paying to serve placeholders, it identifies which control users would have won the advertiser's ad, producing a matched counterfactual at negligible cost. Their implementation recorded more than 100 million predicted ghost ads a day.
Productisation followed, and standards bodies arrived late. The IAB and IAB Europe published an incrementality framework for commerce media on September 9, 2025, followed by Guidelines for Incremental Measurement in Commerce Media on November 3, 2025. Access widened when Google cut its minimum experiment budget to $5,000 on November 11, 2025, a threshold first announced at Google Marketing Live in May 2025 alongside the shift to Bayesian estimation.
Why the design matters commercially
The IAB guidelines rank methods by causal strength, and experiment-based approaches including holdouts sit at the top, ahead of modelled counterfactuals, econometric models and platform-reported metrics. That hierarchy is now embedded in how budgets get defended. Meta's 'suite of truth' framework, published on May 28, 2025, places lift studies above attributed sales, and retail media networks have adopted geo-holdouts and control groups to make results comparable across networks. Holdout results increasingly serve as calibration priors for marketing mix models rather than standalone verdicts, the role Google built Meridian GeoX to fill when it announced the tool on May 5, 2026.
Limitations and disputes
Precision remains the binding constraint, and a small control group compounds it. Confidence figures describe the probability that a measured gap is not chance. They say nothing about the size of the effect or how the holdout was sized, both of which govern how far a result generalises.
Contamination is the second problem. A control user reached by the same advertiser on an unmeasured platform, or a consumer crossing a regional boundary to buy, compresses the measured gap and understates lift. Google's geographic documentation treats this as a design parameter rather than an edge case.
Error in the identity layer beneath an experiment can swamp the experiment itself. Research published by LiveRamp and the Marketing and Media Alliance on July 20, 2026 simulated a campaign with a true 25% lift and found that 50% identity precision read it as roughly 6.8%, with modelled return falling from $1.50 to $0.43. That is an illustrative lower bound from a vendor selling identity infrastructure, but the mechanism is not disputed.
Conflict of interest sits close to the surface, since the party running the experiment is frequently the party selling the media. Academic work on divergent delivery, the tendency of delivery algorithms to route campaign variants to different audiences, complicates that critique rather than confirming it. A study of Meta's tools covering 3,204 Lift tests found 0.16% of standardised mean differences above the conventional 0.2 threshold, against 22% across 181,890 A/B tests. Three of its five authors are Meta employees, and they argue the problem is specific to A/B tests. If that holds, criticisms of platform A/B testing cannot be transplanted onto holdout designs.
A newer objection targets how results are consumed. A preprint filed on August 21, 2026 argues that collapsing an experiment into a single lift figure for use as a mix-model prior discards the temporal structure that would recover adstock and saturation directly.
Adjacent terms
An A/B test compares two campaign configurations without a no-ad control, so it measures relative performance rather than incremental effect. Ghost ads and ghost bidding are holdout variants that identify or bid for control impressions without serving them, cutting the cost of the counterfactual. A geo experiment randomises at market level rather than user level, trading identity precision for coverage of offline outcomes. Brand lift measures survey-reported perception rather than behaviour. Always-on incrementality reads natural variation in spend rather than imposing a control group at all.
Recent developments
Vendors spent 2026 attaching holdout logic to buying platforms. Nexxen built ghost bidding into its demand-side platformon March 24, 2026. Smartly signed a letter of intent to acquire INCRMNTAL on March 16, 2026. Innovid added purchase-impact and control-group measurement on April 29, 2026, and Kochava folded self-serve on/off pulse testing into its core product during the second quarter.
Connected television has been the most active surface. Jamloop opened a household-level holdout methodology on July 21, 2026, reporting 3,224 incremental subscriptions at 99.98% confidence and an estimated 365% return for one streaming client, a vendor-supplied figure covering a single campaign. Lifesight published cross-channel lift figures in August 2026 claiming a 22.3% paid search conversion lift when campaigns ran alongside streaming.
Reporting infrastructure caught up last. Google Ads API v25.1 exposed 24 Conversion Lift metrics and two read-only resources covering study configuration and flight dates in August 2026, with no ability to create or terminate a study programmatically. The Trade Desk listed conversion lift enhancements in closed beta in its Zuma release on August 27, 2026, and LoopMe brand lift studies opened to its buyers at a one million impression minimum in early September 2026.
Timeline
- December 2011: Jon Vaver and Jim Koehler publish Google's geo experiments methodology, establishing market-level randomisation as a measurement design
- June 2015: The ghost ads working paper circulates, proposing a low-cost counterfactual for control groups
- November 2015: Lewis and Rao publish their analysis of twenty-five field experiments in the Quarterly Journal of Economics
- December 2017: Ghost ads appears in the Journal of Marketing Research and later receives the Paul E. Green Award
- May 22, 2025: Google announces a $5,000 minimum experiment budget and Bayesian estimation at Google Marketing Live
- May 28, 2025: Meta publishes its hybrid measurement framework positioning lift studies as the reference method
- September 9, 2025: IAB and IAB Europe publish an incrementality framework for commerce media
- November 3, 2025: IAB and IAB Europe publish Guidelines for Incremental Measurement in Commerce Media
- November 11, 2025: Google lowers its incrementality testing threshold and adds feasibility ratings
- March 16, 2026: Smartly signs a letter of intent to acquire INCRMNTAL
- March 24, 2026: Nexxen adds ghost bidding incrementality to its demand-side platform
- April 29, 2026: Innovid adds purchase-impact and control-group measurement
- May 5, 2026: Google announces Meridian GeoX, an open-source geographic experimentation tool
- July 20, 2026: LiveRamp and the Marketing and Media Alliance publish identity precision simulations
- July 21, 2026: Jamloop opens household-level holdouts in its self-serve platform
- August 21, 2026: A preprint argues that collapsing geo-experiments into single lift figures discards dynamics
- August 2026: Google Ads API v25.1 exposes 24 Conversion Lift metrics and read-only study resources
Related PPC Land coverage
- Explaining incrementality - The broader concept holdout studies estimate, including the method families the IAB grades by causal strength.
- Explaining endogeneity - Why observational estimates inflate when targeting concentrates exposure on likely converters, and what randomisation fixes.
- Google lowers incrementality testing threshold to $5,000 for advertisers - The November 2025 change, feasibility ratings and campaign type eligibility.
- Google cuts incrementality testing budget requirements to $5,000 minimum - The original Google Marketing Live 2025 announcement and the shift to Bayesian methodology.
- Google Ads API v25.1 gains 24 lift metrics, allowlist only - Read-only programmatic access to study configuration, flight dates and lift metrics.
- Meta's 'suite of truth' framework rewrites how advertisers measure ad impact - The platform position placing lift studies above attributed sales.
- IAB unveils incrementality framework for commerce media budgets - The September 2025 framework distinguishing incrementality from attribution and correlation.
- IAB releases measurement framework for commerce media campaigns - The November 2025 guidelines and the standards work around them.
- How retailers are finally solving the audience targeting puzzle - Geo-holdouts and control groups inside retail media measurement principles.
- LiveRamp study: identity errors cut campaign ROI 70%, killing profit - Simulations showing a true 25% lift read as 6.8% under poor identity precision.
- Nexxen bets on unified AI optimization to fix CTV's measurement gap - Ghost bidding implemented as a synthetic holdout inside a live auction.
- Smartly signs LOI to buy INCRMNTAL, adding always-on incrementality - The continuous alternative that avoids holdout groups and paused media.
- Kochava's Q2 bulletin adds 19 partners, drops agency need for testing - Self-serve on/off pulse testing folded into a core measurement product.
- Innovid expands measurement with purchase data and publisher attribution - Purchase-impact and control-group measurement added to a third-party stack.
- Jamloop measures 3,224 incremental subscriptions at 365% ROAS on CTV - Household-level holdouts suppressed before the auction, and the sizing questions the headline figure leaves open.
- CTV lifts paid search conversions 22.3%, Lifesight report finds - Cross-channel lift claims and the vendor landscape around them.
- MMM overstates paid search ROAS by 2.5 times, Zalando researcher finds - The argument that summarising experiments into single priors discards information.
- Meridian lands inside Analytics 360 as Google links ad spend to future sales - Meridian GeoX and the integration of experiment results into mix modelling.
- Trade Desk says Koa AI in Zuma cuts CPAs 32% across 62 campaigns - The August 2026 Kokai release listing conversion lift enhancements in closed beta.
- Trade Desk buyers gain LoopMe brand lift studies at 1 million impressions - Impression thresholds that determine which campaigns can be measured at all.
Summary
Who. Advertisers and agency measurement teams commission holdout studies; platforms including Google, Meta, Nexxen and The Trade Desk run them inside their buying interfaces; measurement vendors such as Kochava, Jamloop, Innovid and Lifesight sell them independently. Standards work sits with the IAB and IAB Europe.
What. A randomised experiment that withholds a campaign from a control group so the difference in outcomes between exposed and withheld groups estimates what the advertising caused. Results are expressed as absolute lift, relative lift, incremental cost per action and incremental return on ad spend.
When. The geographic form was formalised at Google in 2011, the statistical limits documented in 2015, and the low-cost ghost ad variant published in 2017. Standardised definitions arrived in September and November 2025, and platform access widened on November 11, 2025.
Where. Inside buying platforms as native lift products, in third-party measurement tools, in clean rooms, and as a calibration input to marketing mix models.
Why. Platform-reported returns credit conversions that would have occurred anyway, and budget moves against those numbers. A holdout study is the only widely available design that produces a counterfactual by construction rather than by assumption, which is why it anchors most current measurement frameworks despite being expensive, imprecise and infrequent.
Discussion