Lookalike modeling is the technique of taking a list of known people, usually an advertiser's customers, and finding a larger group of strangers who resemble them. The advertiser supplies the list. The platform matches it to its own users, learns what those users have in common, scores everyone else on the same traits and returns the closest matches as a targetable audience. The method exists because the most valuable customers are already known, while the people most likely to become customers are not. Lookalike modeling bridges the two without the advertiser ever having to describe, in demographic or interest terms, who it wants to reach.

The output goes by many names. Meta calls it a Lookalike Audience, Google has used both "similar audiences" and "Lookalike segments", Amazon DSP calls it Similar Audiences and Pinterest calls it an actalike.

How a lookalike is built

The process starts with the seed, also called the source audience. It can be a hashed customer file, a list of website or app visitors, people who watched a brand's videos, or buyers ranked by value. The platform first matches the seed to accounts it can recognise. Only matched people count, so a file of 10,000 email addresses may yield a seed of a few thousand.

Platforms set minimums. Meta requires a source of at least 100 people from a single country and recommends 1,000 to 5,000, according to a guide by Jon Loomer, a consultant who documents Meta's advertising tools. Google's Lookalike segments in Demand Gen needed more than 100 active matched users across the combined seed lists, and Pinterest asks for 100 matched members.

The second step is representation. Every user is described as a long vector of signals: pages followed, videos watched, purchases, searches, device, location and time of activity. The model infers which of them matter.

The third step is scoring. Two broad families of method exist. Similarity approaches compare every candidate with the seed directly, using nearest-neighbour search or graph links between users. Model-based approaches train a classifier that treats seed members as positive examples and a sample of the general population as unlabelled ones, then predict for every user the probability of belonging to the seed. A 2016 paper by Qiang Ma and colleagues at Yahoo described a graph-based system with nearest-neighbour filtering serving thousands of campaigns, and reported that it beat three other lookalike systems by more than 50% in conversion rate for app-install campaigns. That figure comes from the system's own developers.

The fourth step is thresholding. Scores are ranked and cut at a share of the eligible population. This is the slider advertisers see. Meta offers 1% to 10% of a country's users; a 1% lookalike in the US is about 2.8 million people, according to Loomer, regardless of how large the seed is. Google's tiers are Narrow at about 2.5% of users in the target location, Balanced at about 5% and Broad at about 10%. Lower percentages buy resemblance; higher ones buy reach.

Last comes maintenance. Most platforms exclude the seed itself from the result and refresh the model on a schedule. Meta's lookalikes update every three to seven days while in active use. Google's refresh every one to two days, and a segment whose seed falls below the minimum can keep serving for up to three days.

From Facebook to every platform

Facebook launched Custom Audiences, which matched hashed customer files against its user base, in autumn 2012. Lookalike Audiences followed on March 19, 2013. At launch there were two settings, according to The Next Web: "Similarity", the top 1% of people in a country who most resembled the custom audience, and "Greater Reach", the top 5%.

Twitter added an equivalent in 2014, Snapchat in 2016 and Pinterest on June 14, 2016. Google had offered "similar users" built from remarketing lists on its Display Network, with a 500-cookie minimum and a 30-day lookback, according to Search Engine Land, and extended similar audiences to Search and Shopping in May 2017. LinkedIn introduced Audience Expansion in 2017 and full lookalikes on March 20, 2019, letting B2B advertisers seed with lists of company names, according to AdExchanger. Independent data companies built their own: Adsquare added lookalike extension for mobile location audiences in January 2019.

Then came a retreat. Google said in late 2022 that it would remove similar audiences from Google Ads and Display & Video 360 (DV360), blocking new use from May 1, 2023 and deleting the segments on August 1, 2023, and pointed to optimised targeting instead. LinkedIn discontinued lookalikes on February 29, 2024, promoting predictive audiences as the replacement. Meta folded "Lookalike expansion" into the Advantage brand in March 2022 as Advantage lookalike, which lets the system go beyond a 1% audience when it predicts better results.

The retreat proved partial. Google kept Lookalike segments in Demand Gen, the campaign type that became generally available in October 2023, and Amazon introduced Similar Audiences in beta across 17 countries in September 2024, seeded with audiences built from Amazon shopping interactions.

Why the marketing community cares

For many advertisers, lookalikes were the main prospecting tool on social platforms for a decade. They turned a customer relationship management (CRM) file into reach at scale, and they did so without third-party cookies, because the modelling happens inside a platform's own logged-in data.

That property now carries weight in clean rooms. In a Decentriq-run deployment, The New York Times Advertising and a wealth manager used the roughly 25% overlap between the bank's customers and Times readers as a seed, built lookalikes without either party seeing the other's raw data, and reported a 25.12% US uplift over control from a 1% audience. Those results were published by the vendor and the publisher, without independent verification. Retail media follows similar logic: Amazon Marketing Cloud's high-value audiences let brands sort customers into spending percentiles and model lookalikes from the top buckets.

When OpenAI added custom audience uploads to ChatGPT ads in May 2026, no lookalike option was visible - a reminder that lookalikes usually follow custom audiences on a new platform.

Discrimination, opacity and accuracy

The most serious criticism is legal. Because a lookalike reproduces whatever patterns the seed contains, a seed skewed by race, sex or age produces a skewed audience even if no protected attribute is ever selected. On March 28, 2019, the US Department of Housing and Urban Development (HUD) charged Facebook with housing discrimination. "Using a computer to limit a person's housing choices can be just as discriminatory as slamming a door in someone's face," said Ben Carson, then HUD secretary.

Earlier that month Facebook had settled suits brought by civil rights groups and created Special Ad Audiences for housing, employment and credit ads, a lookalike variant with demographic inputs removed. Researchers led by Piotr Sapiezynski at Northeastern University tested it. A Special Ad Audience seeded only with women still delivered to 91.2% women, against 96.1% for a standard lookalike. One seeded with Facebook employees delivered to 88% men, against 54% for a random baseline. The authors concluded that merely removing demographic features from inputs "can fail to prevent biased outputs."

On June 21, 2022, the US Department of Justice sued Meta under the Fair Housing Act, alleging that the Special Ad Audience tool, "previously called 'Lookalike Audience'", considered protected characteristics. The settlement required Meta to stop using it by December 31, 2022 and pay a civil penalty of $115,054, the maximum available. Meta withdrew Special Ad Audiences for all special ad categories between August 25 and October 12, 2022, and built the Variance Reduction System, which adjusts delivery when the audience seeing a housing ad drifts from the eligible population. Google applies its own limits: advertisers in 21 sensitive interest categories cannot use lookalike segments at all.

Opacity is the second complaint. Platforms do not disclose which signals drive similarity, so inclusion cannot be audited. Accuracy is the third. A lookalike predicts resemblance, not intent, and a small or noisy seed teaches the model the wrong traits. Reported performance gains almost always come from the platforms or their vendors.

Not the same as

  • Audience overlap analysis measures how far two existing audiences share members. It describes audiences; lookalike modeling creates new ones.
  • Suppression audiences are lists of people deliberately excluded from a campaign, such as recent buyers. A seed list is often reused as a suppression list, but the purposes are opposite.
  • Predictive audiences score known users on the likelihood of a future action, such as Google Analytics 4's "likely 7-day purchasers". They forecast behaviour; a lookalike forecasts resemblance to a seed.
  • Optimised targeting and audience expansion use an advertiser's inputs as hints and follow conversion signals wherever they lead, without a fixed similarity threshold. Unlike cohort-based targeting, all of these resolve to individual users.

Recent developments

On February 17, 2026, Google said Demand Gen Lookalike segments would move in March from hard similarity thresholds to "audience suggestions", letting its systems reach beyond the selected tier. Meta's move to the Advantage+ campaign structure treats lookalikes as one of several inputs that its system may relax.

Elsewhere the explicit control is returning. As part of an overhaul of DV360, lookalike audiences replaced legacy audience expansion for YouTube reach and view campaigns from June 15, 2026, offering 2.5%, 5% and 10% tiers against expansion's previous 1% cap. And on September 1, 2026, Google documented Lookalike segments for Video brand campaigns, where the tiers still act as a boundary. As of October 2026, the same Google feature therefore behaves as a hard limit in one campaign type and a suggestion in another.

Timeline

  • Autumn 2012 - Facebook launches Custom Audiences for hashed customer files.
  • March 19, 2013 - Facebook launches Lookalike Audiences with 1% "Similarity" and 5% "Greater Reach" settings.
  • 2014 - Twitter adds lookalike targeting.
  • 2016 - Snapchat follows; Pinterest launches lookalike targeting on June 14, later renamed actalike.
  • 2016 - Yahoo researchers publish a graph-based lookalike system at a KDD workshop.
  • May 2017 - Google extends similar audiences to Search and Shopping campaigns.
  • 2017 - LinkedIn launches Audience Expansion.
  • January 2019 - Adsquare introduces lookalike audience extension for mobile.
  • March 2019 - Facebook settles civil rights suits and introduces Special Ad Audiences.
  • March 20, 2019 - LinkedIn rolls out lookalike audiences.
  • March 28, 2019 - HUD charges Facebook with housing discrimination.
  • December 16, 2019 - Sapiezynski and colleagues publish measurements of bias in Lookalike and Special Ad Audiences.
  • March 2022 - Meta renames lookalike expansion as Advantage lookalike.
  • June 21, 2022 - The Department of Justice sues Meta and announces a settlement covering Special Ad Audiences.
  • August 25 to October 12, 2022 - Meta phases out Special Ad Audiences.
  • January 2023 - Meta announces the Variance Reduction System.
  • May 1, 2023 - Google blocks new use of similar audiences.
  • August 1, 2023 - Google removes similar audiences from all campaigns.
  • October 2023 - Demand Gen becomes generally available with Lookalike segments.
  • February 29, 2024 - LinkedIn discontinues lookalike audiences.
  • May 2024 - Amazon Marketing Cloud introduces high-value audiences with lookalike options.
  • September 2024 - Amazon DSP launches Similar Audiences in beta in 17 countries.
  • March 2026 - Demand Gen Lookalike segments become audience suggestions.
  • June 3, 2026 - Google clarifies that sensitive categories cannot use lookalike segments in Demand Gen and Discovery.
  • June 15, 2026 - DV360 lookalikes replace audience expansion for YouTube Video campaigns.
  • September 1, 2026 - Google documents Lookalike segments for Video brand campaigns.

Summary

Who. Lookalike models are run by advertising platforms, including Meta, Google, Amazon, Pinterest, Snap and X, and by DSPs, data companies and clean-room providers. Advertisers supply the seeds. Regulators, notably HUD and the Department of Justice in the US, have set limits on their use for housing, employment and credit ads.

What. It is the technique of expanding a known seed audience into a larger one by scoring a platform's users on their resemblance to the seed, then keeping the closest 1% to 10% of a market. Methods range from nearest-neighbour similarity to classifiers trained on seed members.

When. Facebook established the category on March 19, 2013. Most major platforms copied it between 2014 and 2019; Google and LinkedIn withdrew their explicit versions in 2023 and 2024, and Google has since reintroduced Lookalike segments in Demand Gen, Video and DV360.

Where. Lookalikes run inside logged-in platforms, in DSPs such as Amazon DSP and DV360, in retail media clean rooms such as Amazon Marketing Cloud, and in independent clean rooms where two parties combine data.

Why. Advertisers know their best customers but not the strangers most like them. Lookalike modeling finds those people without third-party cookies or hand-built targeting, at the cost of opacity, uneven accuracy and a documented tendency to reproduce the demographic skews of the seed.