Reddit chief executive Steve Huffman used a June 22, 2026 podcast recording, released July 3, to put a number on how much of modern language models rests on his platform, and to explain why the company that once gave its data away now litigates over it.

The interview appeared on Mixed Signals, the media podcast produced by Semafor, and was published on YouTube on July 3, 2026. Media editor Max Tani and editor-in-chief Ben Smith recorded it at Cannes Lions, on the day Reddit turned 21. Huffman co-founded the site in 2005 and launched it on June 22 of that year, at the age of 21. "So today is Reddit's 21st birthday," he said, noting that the platform now accounts for exactly half his life.

Much of the conversation concerned a question that has become commercially significant for anyone buying media against user-generated content: how much of what large language models produce originates on Reddit, and what that content is worth.

Putting a figure on the training set

Asked how much of the output from Gemini, Claude or ChatGPT traces back to Reddit, Huffman was blunt. "A lot. We don't know for sure and they don't tell us, but it's a lot," he said.

He then supplied the only public anchor he has. "One data point I have is OpenAI's last public research paper," he said, adding that it "said that Reddit was about a third of their training set." He placed that paper at the GPT-2 or GPT-3 stage and acknowledged that nothing comparable has been disclosed since.

Huffman attributed the concentration to the way people write on the platform. Conversations there capture natural speech rather than the register of a book or an encyclopedia entry, and the subject range spans hobbies, purchases, medical worries and current events. Smith summarised the consequence more sharply: "I think the language of LLMs is largely the language of Reddit."

The company did not plan for that outcome, according to Huffman. "It was neither intentional nor an accident," he said, before drawing a distinction between his own position and that of the model developers. He described the use of Reddit content without permission as something no frontier laboratory avoided, and characterised the shift from academic research to commercial deployment as the point at which the arrangement broke down. Reddit still supplies data to universities, he said, and the company was historically not precious about it.

His formulation of the current stance was compact: "commercial use requires commercial terms." Where a project is run for public benefit rather than profit, he said Reddit remains open to supporting it. What has changed, in his account, is the emergence of what he called an industry around laundering data, and a conviction that "people are taking it, using it to enrich themselves and lying about it."

Those remarks sit on top of an existing set of commercial and legal arrangements. Reddit signed a data licensing agreement with Google in February 2024, reported at roughly 60 million dollars a year, and followed it with a partnership with OpenAI announced on May 17, 2024. In July 2024, changes to the site's robots.txt file effectively made Google the only search engine able to index recent Reddit content, a move that drew criticism over search competition and AI training ethics.

The adversarial track began on June 4, 2025, when Reddit filed a 28-page complaint against Anthropic in San Francisco Superior Court, alleging breach of contract, unjust enrichment, trespass to chattels, tortious interference and unfair competition. The filing claimed Anthropic crawled the platform more than 100,000 times after stating publicly that it had stopped, a contention that remains a reference point in debates over crawler compliance.

Huffman said he would prefer relationships with every model developer rather than an exclusive arrangement, on both business and structural grounds. "We don't want one AI winner," he said. "We don't want one search winner." He argued that the underlying technology is powerful enough, and understood poorly enough, that the ecosystem needs more participants rather than fewer.

The economics of that position have drawn scrutiny. During the company's first-quarter 2026 earnings call, an analyst pressed Huffman on whether licensing terms reflect the value of the underlying data, and he declined to discuss specific deal economics.

Brand safety moved from objection to selling point

Reddit's advertising business rests on ground that looked very different a decade ago. Huffman said that when the company met marketers ten years ago, "the first and last question they'd ask about is brand safety." The platform then had no safety team and, by his description, close to no content policy, a posture he traced to a founding assumption that consenting adults should decide their own subject matter.

That changed through policy work, a dedicated safety function and enforcement practice. "Now, it's a topic that very rarely comes up," he said, describing the subject as something closer to an advantage than an obstacle in commercial conversations. Brand safety has meanwhile hardened into a measurable procurement category across the wider market.

Audience composition shifted alongside it. Huffman said the early user base was "basically all 20some dudes," and attributed part of that narrowness to an interface poor enough that only people accustomed to bad interfaces, principally gamers, tolerated it. Every improvement in usability, and every reduction in the odds that a new visitor would open the app and conclude it was not for them, produced incremental growth. His stated growth strategy across two years of earnings calls has been to "make Reddit better."

The financial trajectory is documented. Reddit reported 726 million dollars in fourth-quarter 2025 revenue with advertising up 75%, then 625 million dollars in first-quarter 2026 advertising revenue, a 74% year-on-year increase, alongside 127 million daily active uniques. Roughly 40% of conversations on the platform are commercial in nature, according to company disclosures. At Cannes Lions on the same day as the interview recording, Reddit introduced four advertising products built on its Community Intelligence layer, including a free-form ad generator and Redditor Highlights.

The slop problem the platform declines to prohibit

Tani asked what Reddit is doing to prevent bot-written material from filling the platform and then re-entering model outputs. Huffman reframed the category. Before AI-generated filler there was spam, and the company treats both as the same operational problem: content produced mechanically that should not be on the site.

On the enforcement side, Reddit has been improving detection of inauthentic behaviour and devaluing spam so it accumulates fewer views, including the karma farming pattern in which an account builds reputation through legitimate posts before pivoting to spam. Huffman's 21st anniversary post on June 16, 2026 disclosed that proactive moderation models block up to 23 million spam views per day, and that the platform revokes close to 2 million inauthentic votes daily.

The more interesting admission concerned the limits of policy. "The most common AI slop on Reddit actually isn't a bot," Huffman said. "It's a normal user who uses AI to write a post and then they paste it into Reddit." Asked whether that is permitted, he said it is, and explained why: "we can't ban that. That's just how people are going to write, but the communities are rejecting it."

His argument is that readers have developed a detection reflex, recognising the characteristic patterns of machine-written prose, and that downvoting supplies the enforcement mechanism a rule could not. Generated content that survives, in his framing, will have to be good enough that people consider it additive.

That thesis runs against a body of measurement in the advertising market. AI slop has become a defined brand safety category with its own vendor tooling, and Integral Ad Science moved its avoidance product to general availability in June 2026. Research from Raptive found that suspected AI-generated content cuts reader trust by nearly half and produces a 14% decline in purchase consideration. Separate analysis has documented how fabricated material recirculates through AI answer engines and returns as apparent fact.

Smith raised Moltbook, a short-lived attempt to build a Reddit for AI agents. Huffman's assessment was that the site read as "clearly a bunch of humans pretending to be bots."

Humanness without identity

Pressed on whether AI will force the platform to abandon pseudonymity, Huffman drew a line between two objectives. "We want to verify humanness," he said, arguing that phones, passkeys and device biometrics such as Face ID or Touch ID already require a person to perform an action, and get most of the way there without exposing a name or a face.

He was dismissive of document-based checks. "The problem with government IDs is is you can literally pay somebody to use their ID and log in for you," he said, adding that such systems mainly curb growth. He said something better than showing a government identity document would be powerful, and referenced decentralised humanness systems of the kind Worldcoin has attempted to build, while noting the difficulty sits in the implementation details.

Anonymity, in his account, is functionally load-bearing. People disclose medical conditions and addiction on Reddit precisely because a real name is not attached, and that candour is what makes the resulting archive useful.

Regulators have taken a different view of what Reddit must know about its users. The UK Information Commissioner's Office fined the company 14.47 million pounds on February 24, 2026 after finding it had no robust age assurance mechanism and therefore no lawful basis to process data belonging to children under 13, and that it had not completed a data protection impact assessment before January 2025. Comparable requirements are being built across the European Union.

A platform designed to prevent influencers

One structural choice separates Reddit from every other consumer platform of comparable scale. "You cannot get rich or famous by being good at Reddit," Huffman said, describing this as deliberate. Reputation does not travel off the platform, which he argued keeps incentives clean.

He framed the company's history in three phases. It was founded as an alternative to traditional media, where editors decide the agenda. It then defined itself against social media by declining to build a social graph, a decision he said cost the company growth. The current phase, he said, is defined against automation: "Reddit is not AI. As the rest of the internet becomes more automated and summarized, Reddit is still messy and human and authentic."

In their closing discussion, Tani and Smith identified the absence of a creator economy as the platform's most unusual asset, and the difficulty of stopping humans from pasting machine-written text as its central vulnerability.

The branded segment and the numbers it cited

The episode carried a branded segment from Think with Google, in which a Google vice president of marketing discussed search volume. The executive said that consumers using advanced search features search more than before, and described a shift in behaviour: "they're kind of transitioning from, I think, queries into quests."

He also made a specific performance claim about the automation suite Google has been pushing into standard search campaigns. "Marketers who turn on AI Max for search are seeing a 7% increase in conversions," he said, describing the product as one of the most successful launches in the company's history.

That figure matches what Google has published elsewhere. When the company announced the retirement of Dynamic Search Ads in April 2026, it cited an average of 7% more conversions or conversion value at similar cost per acquisition when the full AI Max feature set runs together, compared with search term matching alone. That is a more conservative number than the 14% uplift cited at the May 2025 launch, and in May 2026 the company quoted 27% more conversions against campaigns built mainly on exact and phrase match.

Independent testing has not reproduced those results consistently. Analysis of more than 250 retail campaigns published in November 2025 found AI Max delivering conversions at approximately 35% lower return on ad spend than traditional match types, with one four-month test recording a cost per conversion of 100.37 dollars under AI Max against 43.97 dollars for phrase match. Earlier work in August 2025 found 99% of impressions across roughly 30,000 AI Max-activated search terms produced no conversions.

Why this matters for the marketing community

Three practical consequences follow from the interview.

First, the training-set figure gives publishers a rare public benchmark. If a single platform supplied roughly a third of one major model's corpus, the negotiating position of any content owner supplying language data is stronger than licensing revenue currently reflects. Reddit's own disclosures show data licensing at 36 million dollars in a quarter against 549 million dollars in advertising in the same period, a ratio that invites the question Huffman declined to answer on the earnings call.

Second, the moderation position sets a boundary that verification vendors cannot cross. Reddit will not prohibit users from writing with AI assistance, which means the inventory advertisers buy on the platform contains machine-assisted text by design, filtered by community voting rather than classification. For buyers applying AI content avoidance across the open web, social environments operate on different logic, and in-app and social inventory has remained outside the coverage of several avoidance products.

Third, the humanness question determines what audience data will exist. A platform that verifies personhood without identity produces different targeting inputs than one holding verified ages and names. Huffman's stated preference forecloses using verified identity as a targeting signal, and regulatory pressure in the United Kingdom and European Union pushes in the opposite direction.

Timeline

Summary

Who: Steve Huffman, co-founder and chief executive of Reddit, interviewed by Semafor media editor Max Tani and editor-in-chief Ben Smith on the Mixed Signals podcast.

What: Huffman said Reddit accounted for about a third of the training set described in OpenAI's last public research paper, set out the company's position that commercial use of its content requires commercial terms, described brand safety as a subject that now rarely arises with advertisers, and stated that Reddit will not ban users from writing posts with AI assistance because community downvoting handles the problem. He also said the company wants to verify humanness rather than identity, using phones, passkeys and device biometrics rather than government documents.

When: The interview was recorded on June 22, 2026, Reddit's 21st birthday, and published on YouTube on July 3, 2026.

Where: Recorded at Cannes Lions and distributed through Semafor Media's Mixed Signals podcast and YouTube channel.

Why: Reddit sits between two businesses that pull in opposite directions. Its advertising revenue depends on human conversation that models consume and summarise elsewhere, while its licensing revenue depends on selling that same conversation to the companies building those models. The interview is the clearest public account to date of how the platform's chief executive reconciles the two.