Karan Dhir, a marketing measurement product lead at Genentech who previously held senior product roles at Amazon and Walmart, argued in a Substack essay published on September 9, 2026 that advertising platforms deliver two incompatible kinds of numbers through the same interface. One kind, he wrote, is causal evidence. The other is operational telemetry, and peer-reviewed research built on Facebook's experimentation infrastructure measured the distance between them years ago.
In Short
A measurement specialist wrote that the ad results platforms show you mix two very different things: real experiments and statistical guesses. Research run on Facebook's data found that the guessing methods often landed far from the experimental answer, frequently several times too high and sometimes too low. What changes is the question asked of every lift number before it moves a budget: did a genuine experiment produce it, or did a model?
A classification problem
The essay carries the title "Facebook Proved Its Own Ad Measurement Wrong" and sits on Dhir's personal Substack, karandhir.substack.com. Promoting it on LinkedIn, he compressed it to two sentences: "This week I wrote about platform lift studies. The argument comes down to a classification discipline."
The subject is how platforms report incrementality, the outcomes a campaign actually causes rather than those that would have happened anyway. Platforms, in Dhir's account, can produce two different classes of measurement of that quantity. "Genuine randomized lift studies count as evidence within their scope. Modeled conversions and algorithmic lift estimates without randomized holdouts count as telemetry," he wrote on LinkedIn. The industry, he added in the essay, absorbed the underlying research "as a vague sense that attribution is imperfect," which he regards as a misreading.
His vantage point is the buy side. According to his LinkedIn profile, Dhir has worked since April 2025 as principal AI product manager at Genentech in South San Francisco, where he describes owning "a decision-grade marketing measurement platform, MMM, incrementality, and ROI attribution" that commercial teams use to reallocate spend across a portfolio of 17 or more brands. In the essay he writes, "I run measurement for a $1B portfolio," a self-reported figure PPC Land has not verified. The piece belongs to a series Dhir later described on LinkedIn as The Incrementality Playbook, a four-week argument.
The research underneath
Two papers in Marketing Science carry the weight of the case. The first, "A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook," appeared in volume 38, issue 2, in March 2019. Its authors were Brett Gordon and Florian Zettelmeyer of Northwestern University's Kellogg School of Management, together with Neha Bhargava and Dan Chapsky, who were at Facebook, according to a release from INFORMS, the journal's publisher. Using 15 US advertising experiments comprising 500 million user-experiment observations and 1.6 billion ad impressions, the team treated each randomized result as ground truth and asked whether observational methods could recover it. The copy of Dhir's essay reviewed by PPC Land gives the observation count as "5 million," apparently truncated; the published abstract reports 500 million.
The methods under test, exact matching, propensity score matching and regression adjustment among them, "weren't straw men," Dhir insists. According to the abstract, they often failed to reproduce the experimental effects even after conditioning on extensive demographic and behavioral variables, and INFORMS summarized the findings as showing that the more common methods overestimated ad effectiveness relative to the randomized tests, though in some cases they significantly underestimated it. Dhir's gloss is harsher. On LinkedIn he wrote that the methods "miss experimental ground truth widely and unpredictably, frequently by a factor of 3 or more, even with rich user-level data." The abstract states no single multiplier, so that figure is his characterization rather than a headline number the authors published.
Why would so much data fail? "The failure is structural," Dhir wrote. "Ad delivery targets people already likely to convert, and no adjustment on observables removes selection operating on unobservables." Economists call the underlying problem endogeneity: exposure to an ad is correlated with demand that no dataset fully captures, so the people who see a campaign differ systematically from those who do not, and the gap inflates apparent effects.
The second paper widened the sample dramatically. "Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement," by Gordon, Robert Moakler and Zettelmeyer, appeared in Marketing Science volume 42, issue 4, in July 2023. It drew on 663 large-scale experiments at Facebook and more than 5,000 user-level features, data the authors described as richer than what most advertisers or their measurement partners can access. Two methods were tested: double/debiased machine learning and stratified propensity score matching.
According to the paper's abstract, median lifts measured by randomization were 29%, 18% and 5% for upper, middle and lower funnel outcomes. The machine-learning approach produced median estimates of 83%, 58% and 24% for the same three outcome types; stratified propensity score matching produced 173%, 176% and 64%. Machine learning beat matching, but neither performed well, and the authors concluded that even with large-scale experiments and rich user-level data they could not reliably estimate a campaign's causal effect. "The 2023 replication removed the remaining excuses," Dhir wrote.
There is a limit to what both papers show. Neither study audited the attributed conversion figures inside Meta's Ads Manager. They benchmarked categories of observational estimator, run on platform data, against randomized outcomes. The step from those estimators to dashboard reporting is Dhir's argument: modeled and attributed figures, he wrote, "come from exactly the observational machinery the Gordon studies benchmarked, or from machinery with even less identification behind it."
Two products in one dashboard
Dhir's first class of measurement is the randomized lift study. Eligible users are assigned at random to a treatment group that sees the campaign or a control group that does not. "When randomization is real, this design inherits the full authority of the randomized controlled trial," he wrote. He credits the ghost ads approach, published by Garrett Johnson, Randall Lewis and Elmar Nubbemeyer in the Journal of Marketing Research in December 2017, with making the holdout studyeconomical: the method identifies the control-group users who would have been served the ad, logs that phantom impression, and serves them something else.
The second class is modeled and attributed measurement: conversion attribution windows, modeled conversions filling privacy gaps, algorithmic lift estimates produced without randomized holdouts, and the estimated performance figures in default reporting. "They're operationally useful for running campaigns. They aren't causal evidence," Dhir wrote.
Modeled conversions are not new. Google first offered them in Google Analytics in 2017 and brought them to Google Ads in July 2019, indicating that modeling would typically raise reported conversions by a low double-digit percentage. Under Google's Consent Mode, conversion modeling becomes available once an account reaches 700 ad clicks over seven days, and practitioner audits have reported recovery rates between 15% and 40%. Attribution windows keep moving too: Meta scheduled the removal of 7-day and 28-day view-through windows from its Ads Insights API for January 12, 2026, then in March 2026 redefined click-through attribution to count only link clicks.
"And the blurring has gotten worse since 2019, not better," Dhir wrote. Privacy changes, identifier deprecation and operating-system consent frameworks removed a large share of observed conversions, and the platforms filled the gap with estimated ones. "Modeled conversions now flow into the same reporting surfaces, and in some configurations into the outcome metrics of studies presented as experiments. Which means an advertiser can be shown a randomized design whose outcome variable is partly synthetic."
That produces what amounts to a third category. The randomization can be genuine while the outcome is estimated, and "the resulting number sits in an uncomfortable middle category the dashboard doesn't label." His response is to extend the questioning to the outcome side: what share of counted conversions were observed, what share were modeled, and whether the read changes if the modeled share is excluded. He says he asks this in vendor reviews. "The answer changes the conversation, and occasionally changes the result."
Where the conflict sits
Both classes, Dhir argued, arrive with similar precision "from the same vendor, and the ambiguity serves the vendor." A platform that sells media and also measures that media carries, in his words, a "conflict of interest that no individual employee's integrity resolves, because the conflict lives in the product design." That design, he wrote, determines which numbers get foregrounded, which studies are easy to run and "which results get framed as performance."
Meta, notably, agrees on the hierarchy. Its May 2025 white paper, "Building a Suite of Truth," ranked randomized experiments above marketing mix models and rules-based attribution on what it called a ladder of incrementality. The same paper, drawing on 307 studies across 54 advertisers, concluded that at the median advertisers undervalued Meta by 31% under rules-based models and would need a multiplier of at least 1.45 to calibrate. Those are vendor-produced figures about the vendor's own media. When Meta revised click attribution in March 2026, it again called incrementality experiments the benchmark for measurement, while promoting an Incremental Attribution setting it said delivered an average 46% increase in incremental conversions.
The question extends to modeling tools. Meta's open-source Robyn marketing mix model prompted a debate in April 2026 over who benefits when a platform builds the model that allocates budget across platforms. A sharper episode came in August 2025, when a former Meta product manager alleged in a London employment tribunal filing that ROAS for Shops ads had been inflated by 17 to 19% through the inclusion of shipping fees and taxes. Those remain allegations; Meta declined to comment at the time of PPC Land's report.
Dhir also points to eBay. The company's paid search experiments, published by Thomas Blake, Chris Nosko and Steven Tadelis in Econometrica in 2015, found that brand keyword advertising produced no measurable short-term benefit. In Dhir's telling, randomized geographic holdouts found near-zero incremental effect for spending "that attributed reporting had valued in the tens of millions."
Three questions
The practical skill, Dhir wrote, "is interrogation." He reduces it to three questions: "Was assignment randomized, and by what mechanism? Who holds the holdout, and can it be audited? Does the outcome metric reconcile to your financials?"
Randomization and mechanism
A genuine lift study, in Dhir's description, can answer the first question concretely: eligible users were split before the auction, the control group was withheld from delivery, and the mechanism is documented. A modeled lift figure cannot, and "the vendor's representative will usually respond with methodology language describing adjustment rather than randomization." He allows no middle ground. "The distinction is binary. The answer is yes with a mechanism, or it's no."
Custody and statistical power
Randomization run, measured and reported entirely inside one platform is, in his phrase, "evidence with a custody problem." Mature arrangements, he wrote, involve third-party measurement partners, exportable experiment-level results, designs the advertiser co-specifies, and power calculations agreed in advance.
Recent product decisions sit uneasily beside that benchmark. Google Ads API version 25.1, released in August 2026, exposed 24 Conversion Lift metrics only to allowlisted accounts, and in read-only form; no developer can create, modify or end a lift study through the API. Entry costs have fallen meanwhile. At Google Marketing Live on May 22, 2025, the company cut the minimum budget for incrementality tests to $5,000 using Bayesian methods. By November 2025, Google was attaching feasibility ratings to proposed studies: those rated High carry a 60 to 90% probability of a conclusive result and those rated Low 0 to 30%, with a recommended minimum duration of 14 days.
Power is where cheap experiments meet Dhir's objection. The economics literature, he wrote, is blunt that "advertising effects are small relative to outcome variance," so "underpowered lift studies produce noise wearing error bars." Run ten underpowered studies for one advertiser, he argued, and one will eventually produce an impressive result by chance. "Study registration and pre-agreed power are the guardrails," he wrote, adding that they exist only when advertisers insist on them.
Reconciliation to the ledger
The third question moves from statistics to accounting. A lift study built on platform-defined conversions "inherits every definitional gap between the platform event and your ledger," Dhir wrote. A lift readout that cannot be walked, "even approximately, to a financial quantity," will win the meeting, he argued, "and lose the CFO." The strongest programs, in his description, feed experiment outcomes into their own measurement stack, "calibrate the MMM against the experimental reads," and reconcile the combined picture to financial actuals, so each platform experiment becomes an "anchored input instead of a free-floating claim."
A preprint posted weeks before the essay points the same way. A paper by Zalando's Niklas Heusch, posted to arXiv on August 21, 2026, found that a conventional marketing mix model reported ROAS of 10.61 against a true 4.20, while a structural method using four geo-experiments recovered 4.14. The demonstration relied entirely on synthetic data with known parameters, and its references include both Gordon papers.
Classification at intake
Dhir's posture toward platform measurement is, in his words, "Not trust. Not boycott." Every platform number entering his organization's decision processes "gets classified at intake as experimental or observational, the way a bank classifies collateral." The categories are written down, and new sources are classified before their numbers circulate. Experimental results with documented randomization, adequate power and reconcilable outcomes are treated as evidence and used to calibrate models. "Everything else gets treated as operational telemetry... useful for running campaigns, inadmissible for judging them."
He frames this as resolving a paradox in the Gordon research. The same company "whose default reporting embodies the observational methods that missed truth by multiples" also built the experimentation infrastructure that established the truth. "Both facts are permanent," he wrote. "The platforms will keep selling modeled confidence because it scales, and they'll keep operating genuine experimentation systems because sophisticated advertisers demand them."
For practitioners, the difficulty is rarely a shortage of numbers. A Kantar survey of 1,935 decision makers, cited in Meta's May 2025 paper, found the average advertiser used 3.8 measurement solutions and 55% had experienced contradictory results across tools. Dhir's framework does not reduce that count; it ranks the inputs. The same tension surfaced in August 2026, when incrementality vendor Measured warned that Google's changes to Target CPA and Target ROAS bidding could shift results in ways stable platform-reported ROAS would not show.
What the essay leaves open
The argument has boundaries its author partly concedes. A platform lift study, in his framing, is the best evidence only within its scope: it measures that platform's causal effect, not the whole marketing system. Some decisions cannot be randomized at all, and his LinkedIn post flagged the next installment: "the moments you can't randomize at all. National launches, sponsorships, brand campaigns that go everywhere at once, and the synthetic control instruments Google and Uber built for exactly those cases."
The research also cuts in both directions. The 2019 paper documented underestimation as well as overestimation, and in the 2023 study the machine-learning estimator came considerably closer to the experimental benchmark than matching did, even if not close enough. The lone commenter on the Substack post, Bruno Gavino of Codedesign.org, wrote on September 15 that every lift study he had run showed most reported performance to be demand that already existed - one practitioner's experience, not a dataset.
Dhir's conclusion leaves little room for ambiguity. "The 2019 study has been public for 6 years," he wrote on LinkedIn. "Organizations still allocating on modeled lift aren't uninformed. They're unwilling." In the essay, he closed on the difference between the two words: "The distance between those two words is the entire distance between a measurement program and a reporting habit."
Timeline
- 2017: Google begins offering modeled conversions in Google Analytics, later extending them to Google Ads in July 2019
- December 2017: Johnson, Lewis and Nubbemeyer publish the ghost ads method in the Journal of Marketing Research
- March 2019: Gordon, Zettelmeyer, Bhargava and Chapsky publish results from 15 Facebook experiments in Marketing Science
- July 2023: Gordon, Moakler and Zettelmeyer publish results from 663 Facebook experiments in Marketing Science
- May 22, 2025: Google cuts the minimum incrementality test budget to $5,000 at Google Marketing Live
- May 2025: Meta publishes its "Suite of Truth" white paper ranking randomized experiments above other methods
- August 20, 2025: A former Meta product manager alleges Shops ads ROAS was inflated by 17 to 19%
- October 2025: Meta schedules the removal of 7-day and 28-day view-through windows from the Ads Insights API for January 12, 2026
- November 2025: Google details feasibility ratings for incrementality experiments
- March 2026: Meta redefines click-through attribution to count only link clicks
- April 2026: Debate over Meta's Robyn raises the question of who benefits from platform-built MMM
- August 11, 2026: Measured warns about Google's Target CPA and Target ROAS change
- August 2026: Google Ads API v25.1 exposes 24 Conversion Lift metrics to allowlisted accounts only
- August 21, 2026: Zalando researcher Niklas Heusch posts a paper finding MMM overstated ROAS 2.5 times on synthetic data
- September 9, 2026: Karan Dhir publishes "Facebook Proved Its Own Ad Measurement Wrong" on Substack
- September 15, 2026: A reader comment on the essay describes lift studies consistently showing reported performance as pre-existing demand
- September 2026: Dhir says on LinkedIn that he has closed The Incrementality Playbook, a four-week series
Related PPC Land coverage
- Meta's 'suite of truth' framework rewrites how advertisers measure ad impact - Meta's May 2025 white paper ranking experiments, MMM and attribution on a ladder of causal rigor.
- MMM overstates paid search ROAS by 2.5 times, Zalando researcher finds - A synthetic-data study showing how geo-experiments can correct a biased marketing mix model.
- Google Ads API v25.1 gains 24 lift metrics, allowlist only - Read-only Conversion Lift data restricted to allowlisted accounts.
- Google cuts incrementality testing budget requirements to $5,000 minimum - The May 2025 Google Marketing Live change to experiment budgets.
- Google lowers incrementality testing threshold to $5,000 for advertisers - November 2025 coverage of feasibility ratings and conclusiveness probabilities.
- Meta's Robyn: who really benefits when a platform builds your MMM? - The conflict-of-interest debate around platform-built marketing mix models.
- Former Meta employee alleges artificial ROAS inflation for Shops ads - Tribunal allegations that shipping fees and taxes inflated reported ROAS.
- Meta rewrites click attribution rules, finally aligning with Google Analytics - Meta's March 2026 split between click-through and engage-through attribution.
- Meta restricts attribution windows and data retention in Ads Insights API - The removal of 7-day and 28-day view-through windows.
- What are Modeled Conversions? - Early coverage of how Google estimates conversions it cannot observe.
- Measured says Google turns $10 CPA targets into instructions on August 17 - An incrementality vendor's analysis of Google's bidding change.
- The attribution illusion - Wharton research on how last-touch attribution can reduce advertiser profits.
- Bid controls shrink as ad measurement breaks: week of August 17 - A weekly roundup of measurement changes across Google, Meta and OpenAI.
- Google's Meridian GeoX exits beta claiming 31% cheaper geo experiments - Google's geo-experiment tool reaching general availability, published the same day as Dhir's essay.
Summary
Who: Karan Dhir, principal AI product manager at Genentech and former senior product manager at Amazon and Walmart, drawing on research by Brett Gordon, Florian Zettelmeyer, Neha Bhargava, Dan Chapsky and Robert Moakler.
What: An essay arguing that ad platforms present randomized lift studies and modeled or attributed figures in the same reporting surfaces, that only the first counts as causal evidence, and that three questions on randomization, custody and financial reconciliation separate them. The essay rests on two Marketing Science papers showing that observational methods, run on 15 and then 663 Facebook experiments, failed to recover experimental results reliably.
When: The essay was published on September 9, 2026, with the underlying papers appearing in March 2019 and July 2023.
Where: Dhir's Substack newsletter and LinkedIn, with the research conducted on US advertising experiments at Facebook.
Why: Privacy changes have pushed more modeled conversions into platform reporting and, in some configurations, into experiment outcomes, while platforms also supply the measurement of their own media. The argument matters to advertisers deciding which platform numbers can justify budget moves and which serve only to run campaigns.
Discussion