A co-founder of super{set} who spent the 2000s selling yield software to web publishers has told businesses that handing proprietary data to large language models repeats the bargain publishers struck with Google two decades ago, and that no reliable technique exists to reverse it.
Tom Chavez published the argument on LinkedIn this week. The captured post carries a relative timestamp of one day rather than an absolute date, and it was reviewed for this article on Saturday, August 8, 2026, when it showed 80 reactions, two comments and three reposts. The text opens with a memory rather than a thesis.
"20 years ago, I walked around Manhattan begging digital publishers not to hand their ad inventory to Google for pennies on the dollar," Chavez wrote.
An argument built on a specific record
The claim carries weight partly because of where Chavez sat during the period he describes. His LinkedIn profile lists him as chief executive and co-founder of Rapt from 1999 to 2008, a San Francisco company acquired by Microsoft and summarised on the profile as "Algorithmic software for pricing, inventory, and yield optimization for leading web destinations." After the acquisition he served as general manager for online publisher businesses at Microsoft from April 2008 to October 2009, then spent eight months as an entrepreneur in residence at Accel Partners from October 2009 to May 2010. He is now co-founder of super{set}.
Rapt sold publishers the tools to price their own inventory. That is the vantage point from which the post reconstructs the decade that followed.
Three eras, one pitch
Chavez organises roughly twenty-five years of internet commerce into three phases. The first is what he calls the pre-Facebook era, when the portals - Yahoo, MSN, AOL - sat alongside the New York Times, Huffington Post and hundreds of smaller publishers. According to the post, he asked those publishers "to forego the nickel today and invest in the ability to sell their advertising for dollars tomorrow."
The verdict he records is blunt. "Nobody listened. They gave it all to Google. From their P&L's, crash landings, and diminished valuations, it's pretty clear how that all worked out," Chavez wrote.
The second phase, in his account, was corrective. Companies drew the lesson that "Data is the coin of the realm; don't cede it to the giants," and the cloud and big-data infrastructure of the 2010s grew out of that understanding.
The third phase is the present one. "Here we are in 2026. The new behemoths are the LLMs, worth trillions, and they're going to companies with a familiar pitch: give us your enterprise data, your knowhow, your crown jewels. Trust us. We're just infrastructure dorks trying to help you make your business better, smarter, stronger. Your data still belongs to you," according to the post. Chavez adds a single line of comparison: "It reminds me of Google's pitch in the early 2000's."
The technical claim beneath the analogy
The post rests on an assertion about machine learning rather than about commercial terms. "One of the hardest, most unsolved problems in machine learning today is how to make a machine forget," Chavez wrote. "No one's solved it, and it's unlikely to be solved any time soon."
He extends that into a description of where the information goes. Once a model has seen enterprise data, according to the post, "it resides deep within a multi-layer model with tunings and weights that no programmer (including all the high temple priests at Anthropic and OpenAI) can see, measure, or understand."
The sentence that drew the most attention in the comment thread states the consequence directly: "You can't selectively lobotomize an LLM. It can never forget what it learned or what it saw, even if the LLMs pretend that it will."
That claim is narrower than a general privacy complaint, and it maps onto a research field with a documented gap between promise and delivery. Academic work reviewed in PPC Land's coverage of a study finding that large language models can themselves qualify as personal data reached a similar conclusion: unlearning techniques are a promising direction, but severe challenges remain, and editing methods modify stored facts rather than excise them. The same limitation surfaced in litigation this month. A complaint against the AI notetaking company Granola, filed over recordings made without disclosed consent, cites the company's own materials as acknowledging that once material has been used in training, isolating or removing specific data from the resulting model is not achievable with current techniques.
What the first analogy actually produced
The historical half of the post can be checked against the record, and the record is unambiguous about direction if not about causation.
Google Network revenue, the line that pays publishers for ads served on properties Google does not own, fell 4% year-over-year to $6.97 billion in the first quarter of 2026, against consolidated Alphabet revenue of $109.9 billion. In the second quarter the segment came in at $7.303 billion, down 1% from $7.354 billion while Google Search advertising rose 17% to $63.3 billion. A year earlier, the same distribution had already reached the point where 90% of Google advertising revenue came from owned properties rather than the network.
Traffic followed the same slope. A randomized field experiment with 1,065 desktop Chrome users, published on the Social Science Research Network, found outbound organic clicks fell 39.8% when an AI Overview appeared and zero-click searches rose 34.5%, while sponsored clicks stayed flat. Bloomberg reporting on publishers whose traffic collapsed after the introduction of AI Overviews documented declines of 70% or more across travel, cooking and lifestyle categories, with several sites closing.
A federal court supplied the legal counterpart. Judge Leonie Brinkema ruled on April 17, 2025 that Google willfully monopolized the publisher ad server and ad exchange markets for open-web display advertising, a finding that describes the mechanics of the dependency Chavez says publishers walked into voluntarily.
The post does not cite any of this. It asserts the outcome from memory and moves on.
The counter-question: who lobotomises whom
Chavez frames the LLM companies as the successors to Google. The more interesting position, from an advertising perspective, is that Google now occupies both seats at once.
The company remains dominant in the surface that matters most. Datos measurements for the first quarter of 2026 put Google at 94.3% of United States search share by March, with AI tools collectively below 2% of desktop visits. Geminiclimbed from 10.41% of US desktop AI users in November 2025 to 16.06% in March 2026, taking share from ChatGPTrather than from search.
Yet the budget picture is less settled. Billy Grace benchmark data published on August 4, 2026 recorded Google's share of tracked advertiser spend falling from 62% to 57%. Similarweb's 2026 landscape report placed ChatGPT advertising penetration at 26% of US desktop chats with a 0.50% click-through rate, an early commercial layer on a surface that did not exist as an ad market three years ago. And the supply side has begun testing its own bargaining position: the Wall Street Journal reported on July 21, 2026 that USA Today, Politico, The Economist, People Inc. and Reuters were all evaluating whether to continue working with Google, with Reddit executives assessing the same question.
Chavez's own thesis, applied consistently, cuts in that direction. If the durable asset is proprietary information rather than distribution, then the intermediary that spent two decades converting publisher content into its own inventory has no structural immunity when a newer layer converts search itself into an answer.
A thread already running through marketing technology
The enterprise-data version of the argument is not novel this month, and it has been made from inside the industry it concerns.
Satya Nadella, chairman and chief executive of Microsoft, published an essay on July 12, 2026 arguing that firms buying AI services pay twice, once in money and again in the proprietary knowledge they must disclose to get useful output. Nadella's mechanism is more granular than Chavez's: not a bulk data transfer but an accumulation he calls exhaust, made of prompts, tool usage and corrections. Where Chavez concludes with a warning, Nadella proposes a trust boundary and five priorities for retaining ownership of an organisation's learning loop.
Contract disputes have tracked the same fault line. HubSpot withdrew updated Contact Discovery terms four days after publishing them, reverting on July 5, 2026 after customers questioned how a setting governing AI model training interacted with a separate enrichment feature. Reddit's advertising terms effective August 7, 2025 committed that its language model providers would not use advertiser materials for training. On the publisher side, Cloudflare replaced its per-crawl pricing model with per-answer compensation on July 1, 2026, arguing that a crawl is a poor proxy for value when one page may be cited in thousands of generated answers.
Each of those developments treats the question Chavez raises as a matter of terms and pricing. His post treats it as a matter of physics, and the distinction matters: a contractual exclusion can be renegotiated, while a model that has already absorbed a training set cannot be partially rewound.
What the post does not establish
The argument has clear limits, and the post acknowledges none of them.
It offers no citation for the assertion that unlearning is unsolved, no distinction between training on customer data and inference over it under a no-retention agreement, and no engagement with the contractual carve-outs that several vendors serving marketing organisations already publish. It does not address enterprise deployments where data is processed without contributing to base model weights, which is the configuration most enterprise agreements now describe. Nor does it quantify the value transferred, either in the publisher era it recalls or the model era it warns about.
The post is an opinion published to a professional network, carrying 80 reactions at review. It names no company, quotes no contract and cites no study.
What it does supply is a testable historical claim from someone who was in the room. "This isn't ancient history. It's near history, two decades, tops. The lesson still holds," Chavez wrote. His closing line is four words: "It's your data. Think twice."
Timeline
- 1999 to 2008: Chavez serves as chief executive and co-founder of Rapt, a San Francisco company selling pricing, inventory and yield optimization software to web publishers, later acquired by Microsoft.
- April 2008 to October 2009: Chavez serves as general manager for online publisher businesses at Microsoft.
- October 2009 to May 2010: Chavez spends eight months as entrepreneur in residence at Accel Partners.
- April 17, 2025: Judge Leonie Brinkema rules that Google monopolized the publisher ad server and ad exchange markets for open-web display advertising.
- July 24, 2025: Research reviewed by PPC Land finds that large language models can qualify as personal data, with unlearning techniques facing severe practical limits.
- August 7, 2025: Google advertising revenue distribution reaches 90% owned properties against 10% network for the first time.
- April 27, 2026: Datos Q1 2026 data shows Google at 94.3% of US search share in March and AI tools below 2% of desktop visits.
- April 29, 2026: Alphabet reports Google Network revenue down 4% to $6.97 billion against $109.9 billion consolidated revenue.
- July 1, 2026: Cloudflare replaces per-crawl charging with per-answer publisher compensation.
- July 4, 2026: A randomized field experiment finds AI Overviews cut outbound organic clicks 39.8% and raise zero-click searches 34.5%.
- July 12, 2026: Satya Nadella publishes an essay arguing that enterprises buying AI services disclose proprietary knowledge to make the models useful.
- July 21, 2026: The Wall Street Journal reports that USA Today, Politico, The Economist, People Inc. and Reuters are evaluating whether to keep working with Google.
- July 22, 2026: Alphabet reports Google Search advertising up 17% to $63.3 billion and Network down 1% to $7.303 billion.
- August 4, 2026: Billy Grace benchmark data records Google's share of tracked advertiser spend falling from 62% to 57%.
- August 8, 2026: Chavez's LinkedIn post is reviewed showing 80 reactions, two comments and three reposts, with a relative timestamp of one day.
Related PPC Land coverage
- Nadella says using AI models forces firms to leak their own know-how: Microsoft's chief executive makes a structurally similar argument about enterprise knowledge transfer, framed around prompts and corrections rather than bulk data handover.
- Study: large language models qualify as personal data: Reviews the research literature on machine unlearning and the gap between erasure rights and what current techniques can deliver.
- Granola sued for recording meetings without consent to train AI models: A complaint citing a vendor's own acknowledgement that data already used in training cannot be isolated or removed afterwards.
- Alphabet Q1 2026: Google Network ad revenue falls 4% as AI reshapes the web: Quantifies the segment that pays third-party publishers, and its sharpest recent quarterly decline.
- Google Search ads gain 17% to $63.3 billion while network drops 1%: The second-quarter counterpart, showing owned-property growth alongside continued network contraction.
- Researchers find Google AI Overviews cut publisher clicks 39.8%: Causal evidence from a randomized experiment on how AI summaries redistribute clicks away from source sites.
- Google's AI search overhaul decimates website traffic: Documents publisher closures and traffic declines of 70% or more following the rollout of AI Overviews.
- Court rules Google monopolized digital ad tech markets: The April 2025 ruling on publisher ad server and ad exchange monopolization, describing the tying mechanics behind publisher dependency.
- Reddit and USA Today face Google exit as search traffic drops 28%: Reports major publishers openly weighing whether to stop supplying Google with content.
- Google's share of ad spend drops from 62% to 57%, Billy Grace finds: Advertiser-level benchmark showing budget share moving away from Google even as search revenue grows.
- ChatGPT loses web share to Gemini and Claude as ad penetration hits 26%: Measures the emerging advertising layer inside AI assistants and the redistribution of chatbot traffic.
- AI still under 2% but growing: Datos Q1 2026 state of search report: Clickstream data placing AI tool usage against Google's continuing search share.
- Cloudflare stops charging AI per crawl and starts paying per answer: Documents the shift from crawl-based to citation-based compensation for content used by AI systems.
- HubSpot kills Contact Discovery terms after customer backlash: A four-day reversal of terms distinguishing AI model training from data enrichment, after customer scrutiny.
Summary
Who: Tom Chavez, co-founder of super{set}, former chief executive and co-founder of Rapt and former general manager for online publisher businesses at Microsoft, addressing businesses considering enterprise data arrangements with large language model providers.
What: A LinkedIn post arguing that the pitch enterprises now hear from LLM companies mirrors the one publishers accepted from Google in the early 2000s, and that the analogy is worse this time because machine unlearning remains unsolved, meaning data absorbed into model weights cannot be selectively removed.
When: The post carries a one-day relative timestamp and was reviewed on Saturday, August 8, 2026, showing 80 reactions, two comments and three reposts. The events it recalls span roughly 2005 to the present.
Where: Published on LinkedIn. The commercial consequences it describes apply to publishers and enterprises globally, with the ad revenue trajectory most visible in Alphabet's quarterly Network segment disclosures.
Why: The post matters to marketers, publishers and platform buyers because it links two dependency questions usually handled separately. The first, whether ceding inventory to a dominant intermediary destroys pricing power, has a documented answer in three consecutive years of Google Network revenue decline and a federal monopolization ruling. The second, whether ceding proprietary data to a model provider is reversible, currently has no technical answer at all, which places the burden entirely on contract terms that several vendors have already revised under customer pressure.
Discussion