What the method is

An A/B test is a controlled experiment in which an audience is randomly divided between a control (version A) and a variant (version B), and the two groups are then compared on a metric chosen in advance. Because both versions run at the same time among comparable people, a gap in results can be attributed to the one element that differs, within the limits of statistical noise. The approach exists because advertising outcomes move for many reasons at once - seasonality, competitor bidding, budget shifts, changes in audience mix - and a before-and-after comparison cannot separate those forces from the effect of a deliberate change.

The element under test can be a headline, an image, a landing page, a bidding strategy, a keyword match type or a budget. Equivalent labels include split test and, in the research literature, a randomised experiment with Control and Treatment groups. Ron Kohavi, Randal Henne and Dan Sommerfield of Microsoft listed these synonyms in a widely cited 2007 paper presented at the KDD conference.

How a test runs

A test begins with a hypothesis and one primary metric. The 2007 paper calls this an overall evaluation criterion: a single number that summarises the objective and, ideally, captures longer-term effects such as repeat visits rather than clicks alone.

The next decision is the unit of randomisation. In Google Ads Search experiments, a cookie-based split assigns each user to either the original or the experiment, so one person sees one version. A search-based split assigns each search separately, so the same person may see both; according to Google's help documentation, that option may reach significant results faster. Display campaigns always use cookie-based splits.

Traffic share comes next. Google's documentation recommends 50% for the most accurate comparison. In Display & Video 360, the audience split is set at creation and cannot be changed once the test starts, and budgets are expected to be proportional to it. Duration follows from conversion volume. Google's Experiments FAQ advises at least four to six weeks, longer when conversions arrive with a delay, and excludes the first seven days of data to allow for ramp-up. Analysis rests on a significance test. Google says it splits data into 20 buckets per arm, applies jackknife resampling to estimate variance and reads a two-tailed 95% confidence interval. Display & Video 360 lets buyers pick a 90% or 95% confidence level and reports a p-value, the probability of seeing a difference at least this large if the two versions in truth performed alike.

The last step is a decision. An experiment can be applied to the original campaign, converted into a new campaign or ended. Google Ads switches on auto-apply for new experiments by default: favourable results are written into the original campaign at the end date unless the advertiser opts out beforehand.

Who runs it

On the buy side, advertisers and agencies configure tests inside the buying platform: Google Ads, Microsoft Advertising, Display & Video 360 and Meta among them. Microsoft Advertising launched Experiments globally on July 22, 2019, defining an experiment as a duplicate of a campaign that gives a controlled environment for a change, and recommending an A/A phase of two weeks in which the original and the copy stay identical.

On the sell side, publishers meet the method mostly as something done to them. In June 2018 Google said AdSense automatic experiments would run on about 10,000 daily impressions each from June 28, replacing a design that used about 5% of traffic. YouTube offers creators a related tool: up to three titles or thumbnails are compared, the winner is chosen by watch time share, and results are labelled Winner, Performed Same or Inconclusive.

Origin and evolution

The statistical groundwork predates advertising technology by a century. William Sealy Gosset, a chemist at Guinness, developed the small-sample t-test in 1908 and published as "Student". Ronald Fisher's agricultural trials at Rothamsted in the 1920s established randomisation as the basis of experimental design. Claude Hopkins described keyed coupons as a way of comparing advertisements in Scientific Advertising in 1923.

Online, the technique spread through large web companies in the early 2000s. A frequently repeated marketing example is Barack Obama's 2008 campaign. Dan Siroker, then its director of analytics and later a co-founder of Optimizely, wrote on November 29, 2010 that a test begun in December 2007 compared four buttons and six media options - 24 combinations - on Google Website Optimizer across 310,382 splash-page visitors. The best combination converted 11.6% of visitors against 8.26% for the original page, a relative gain of 40.6%. Siroker's extrapolation to roughly 2.88 million extra sign-ups and $60 million in donations rests on assumptions about volunteer conversion and donations per address, and has not been independently verified.

Tooling then moved inside the ad platforms. According to Search Engine Land, Google retired AdWords Campaign Experiments in favour of Drafts & Experiments from October 2016, with older experiments allowed to run until February 1, 2017. Google's free Optimize website tool left beta in March 2017; Google announced in January 2023 that Optimize would close after September 30, 2023. In January 2022, a new Experiments page let advertisers create a custom experiment for a campaign without first building a draft.

New campaign types then received their own formats. Performance Max Experiments arrived in January 2023, comparing the format with existing campaigns. In October 2023 came uplift experiments, which measure what Performance Max adds alongside Search, Video, Discovery and Display campaigns, and Demand Gen experiments, limited to two arms. Asset tests for retail Performance Max campaigns appeared in October 2024 and reached beta for every Performance Max campaign type in January 2026. Single-campaign broad match experiments in June 2025 ended the need to duplicate a campaign for such comparisons.

Why it matters

Automation has enlarged the set of decisions that advertisers cannot inspect directly. Automated bidding, Performance Max and AI Max settings determine where ads run, and an experiment is the main sanctioned way to compare one configuration with another on the same budget. In Google Ads, outcomes also flow straight back into live campaigns under the auto-apply default, so the design of a test affects spending as well as learning.

Cost is the other consideration. Each test routes part of the traffic to an option that may underperform, and an underpowered test can end inconclusive after consuming budget.

Limitations and open disputes

Peeking is the best-documented statistical problem. Ramesh Johari, Leo Pekelis and David Walsh showed, in a paper first posted on December 15, 2015, that standard p-values become unreliable when experimenters stop a test after repeatedly checking the results, and proposed "always valid" measures that hold whenever the test is stopped. Remi Kerhoas of Eskimoz wrote in Search Engine Land that stopping at the first sign of significance, blending mobile with desktop, or ignoring seasonality can distort PPC results.

Divergent delivery is the most contested platform issue. Where a platform compares ad variants without a no-ad control, its delivery algorithm may show each variant to different audience segments, so a difference can reflect who saw an ad as well as what it said. Gordon Burtch, Robert Moakler, Brett Gordon, Poppy Zhang and Shawndra Hill, in a paper submitted on August 28, 2025, compared 3,204 Meta Lift tests with 181,890 Meta A/B tests and found clear audience imbalance in the latter but no meaningful imbalance in the former. Configuration choices reduced the problem without removing it. PPC Land's holdout study explainer notes that three of the five authors work at Meta and treats the claim as one to weigh rather than settled fact.

The same article makes a further point: a comparison of two versions measures relative performance, not incremental effect, because neither arm is denied advertising. It also cites an analysis by Lewis and Rao of 25 field experiments covering $2.8 million of ad spend, in which median ROI confidence intervals exceeded 100 percentage points - a reminder that noisy outcomes demand large samples.

Opacity is a third complaint. Google says advertisers cannot pool several experiments and recompute the statistics afterwards, because they lack the user-level data needed to rebuild the buckets.

Not the same as

  • Multivariate test. Changes several elements at once and measures combinations. The Obama test was a full-factorial design with 24 combinations, which left about 13,000 visitors per cell, so the format needs far more traffic than a two-way comparison.
  • A/A test. Gives both arms the same experience to validate the set-up. At 95% confidence, a sound system should flag a false difference about 5% of the time. Google expects no significant gap in clicks, impressions, CTR or CPC.
  • Holdout or lift test. Withholds the campaign from a control group to measure what advertising caused. Google permits holdbacks of 1% to 50% of eligible traffic in user-based Conversion Lift. Explaining incrementality covers this causal question.

Recent developments

In January 2026, asset tests for Performance Max let advertisers set unequal splits such as 80/20 or 50/50 and lock the asset group while a test runs. On August 4, 2025, Search Engine Land reported that AI Max experiments run inside the original campaign with a 50/50 budget split.

On August 20, 2026, Google described a test that compares budgets or ROI targets across several Search campaigns, with rollout starting in September 2026. PPC Land's report on the announcement noted that the post gave no performance figures and that the rollout may overlap with a bidding recalibration and the AI Max migration, which would make results hard to attribute. Separately, Google lowered the minimum budget for incrementality experiments from levels approaching $100,000 to $5,000 on November 11, 2025, extending causal testing to smaller advertisers.

Timeline

  • 1908: William Sealy Gosset develops the small-sample t-test at Guinness and publishes as "Student".
  • 1920s: Ronald Fisher's field trials at Rothamsted establish randomisation in experimental design.
  • 1923: Claude Hopkins publishes Scientific Advertising, describing keyed coupons for comparing ads.
  • August 2007: Kohavi, Henne and Sommerfield present their controlled-experiments paper at KDD.
  • December 2007: Test begins on Barack Obama's campaign splash page, using Google Website Optimizer.
  • November 29, 2010: Dan Siroker publishes his account of the Obama test.
  • October 2016: Google begins retiring AdWords Campaign Experiments in favour of Drafts & Experiments; older experiments run until February 1, 2017.
  • March 30, 2017: Google Optimize and Optimize 360 leave beta.
  • June 28, 2018: AdSense automatic experiments move to roughly 10,000 daily impressions each.
  • July 22, 2019: Microsoft Advertising launches Experiments globally.
  • January 2022: Google Ads introduces a new Experiments page with single-step custom experiments.
  • January 20, 2023: Google introduces Performance Max Experiments.
  • September 30, 2023: Google Optimize closes.
  • October 2023: Google introduces Performance Max uplift experiments and Demand Gen experiments.
  • October 2024: Asset experiments arrive for retail Performance Max campaigns.
  • June 30, 2025: Google Ads runs broad match keyword tests inside a single campaign.
  • August 4, 2025: AI Max experiments reported.
  • August 28, 2025: Meta divergent delivery paper submitted to arXiv.
  • January 2026: Asset experiments open in beta to all Performance Max campaign types.
  • August 20, 2026: Google announces budget and ROI target tests across several Search campaigns, with rollout from September 2026.

Summary

  • Who: Advertisers, agencies, publishers and creators run the tests; the buying platforms (Google, Microsoft, Meta and others) randomise the audience and calculate the statistics. Researchers at Microsoft, Optimizely and Meta have shaped the practice.
  • What: A randomised comparison of two versions of one element, read on a pre-chosen metric, to isolate the effect of a change.
  • When: Statistical foundations date from 1908 to the 1920s, advertising use from 1923, web-scale use from the early 2000s, and ad-platform experiment tools from 2016 onwards, with new formats added through 2026.
  • Where: Inside ad platforms, on landing pages and websites, in email, in publisher ad units and on video platforms.
  • Why: Attribution and before-and-after comparisons cannot separate a change's effect from noise. Concurrent, randomised arms can, though divergent delivery, peeking and the lack of a no-ad control limit what the result proves.