Three days of trade reporting produced one recurring problem, stated four different ways. Something that used to carry a price has stopped carrying one, or has started carrying one for the first time, and nobody involved agrees on the number.
In New York, the chief executive of a newspaper company put a figure on the annual cost of producing journalism and set it beside the compute budgets of the firms that read it. In Washington, three members of the House introduced a bill that would fine anonymous crawlers $53,000 per violation. In Mountain View, a report that counts how often a publisher's pages appear inside generative answers reached every account holder, still without click data attached. In San Bruno, the video platform doubled the threshold a creator must clear before advertising revenue starts flowing. In the middle of the programmatic supply chain, a brand that once defined an entire product category was retired into a parent company, because the category had dissolved into plumbing.
None of these developments referenced the others. Taken together across August 10 to August 12, 2026, they describe a market in which the unit of value keeps moving and the meter keeps running.
The cost of half a million works
The most useful number of the week came from an interview rather than a filing.
Meredith Kopit Levien, president and chief executive of The New York Times Company, told Bloomberg's Odd Lots podcast that her company spent close to $2 billion in the prior year producing about half a million works of journalism. The interview was published on August 10, 2026 and recorded the day after the publisher's second quarter results, ran for just over an hour, and covered litigation, licensing, newsroom machine learning, video economics, subscription pricing and acquisition policy. Hosts Joe Weisenthal and Tracy Alloway conducted it.
The figure matters because of what has been missing from the licensing debate. Negotiations between publishers and model developers have run for close to three years with almost no public data on the supply side. Deal values leak occasionally. Production costs almost never do. A stated annual outlay of close to $2 billion for roughly 500,000 stories, photographs and videos yields a per-unit reference point of around $4,000 that other publishers, and the buyers across the table from them, can now argue about openly.
The archive behind that annual output runs to 175 years.
Three tests, described in sequence
Kopit Levien set out a licensing structure with three components, and the sequence is what distinguishes it from the usual publisher complaint.
The first component is continuity with the publisher's own commercial trajectory. A workable agreement, in her formulation, is one that can run alongside a strategy to build a sustainable and growing business for a publisher or anybody producing high quality information. That test rules out deals that generate a one-off cash injection while cannibalising the subscription funnel underneath.
The second is permission with control attached. The publisher grants clear permission for its work to be used and retains control over the ways it may be used. This is a governance question about downstream deployment, not a pricing question, and it is the component that a blanket corpus licence fails most obviously. An unrestricted grant for a flat annual sum satisfies neither the continuity test nor the control test, whatever headline number sits on top of it.
The third is price. There has to be a sustainable and fair value exchange for the use of the work.
"The times is very open to dealing when we can find those terms," Kopit Levien said.
Separating those three is the analytical contribution. Publisher negotiations routinely conflate permission with payment, treating a licensing agreement as a single transaction rather than a governance framework with a price attached. Distinguishing them changes what a negotiation can produce. A publisher that has settled control terms can price them. A publisher that has not settled control terms is pricing an unbounded liability.
What a deal built on those tests would contain
The three tests are worth pushing past their stated form, because each one implies contractual machinery that does not currently exist in most publisher agreements.
Continuity implies a carve-out structure. A publisher whose business depends on converting sampling readers into paying subscribers cannot license a corpus in a way that lets a model answer the questions that drive conversion. That means either temporal restrictions, holding recent output back from the licensed set for a defined window, or surface restrictions, permitting training but not retrieval-augmented generation against live content, or both. Neither is exotic. Both are difficult to price, because the value to the licensee falls sharply as the restrictions tighten, and the licensee knows it.
Control is harder, and the reason is technical rather than commercial. Enterprise data cannot be removed from a large language model once it has been trained on it, as a super{set} co-founder argued at the start of the same week, drawing the parallel with publishers handing inventory to Google two decades ago. If a control term cannot be unwound after breach, then control is not enforceable through remedy after the fact; it has to be enforced through restriction at the point of ingestion, and audited by somebody. There is no established audit standard for what a model retained. A publisher granting control rights it cannot verify is granting a term that exists only on paper.
Price then has to absorb the residual risk from both. That is the mechanism by which an input cost figure becomes negotiating leverage rather than a talking point. If the control terms cannot be verified, the price has to compensate for unverifiable control, and the reference point for that compensation is what the input cost to produce.
The market has exactly one public benchmark for the downside case. Anthropic's $1.5 billion settlement in the authors' matter is the largest publicly reported figure attached to a copyright claim in the sector, and it functions as a rough ceiling estimate for what unlicensed use has cost when litigated to resolution. A publisher pricing a licence against that number, rather than against a hypothetical, is doing arithmetic the counterparty can follow.
There is also a supply argument that has strengthened considerably in the past year. The corpus these systems were trained on is not replenishing at its historical rate. Stack Overflow recorded 1,442 questions in July, down 99 percent from a March 2014 peak of 207,204 in a single month, with the whole of 2025 producing 108,981 questions, fewer than one peak month a decade earlier. The corpus that trained the models is being consumed by the products built on it. Where fresh, verified, human-produced material continues to be generated at scale and at cost, its scarcity value rises rather than falls.
Concentration on the buyer side cuts the other way. A bipartisan amicus brief filed on August 4 in the D.C. Circuit put the point on the record, arguing in 34 pages that three firms hold 88 percent of AI model API revenue. A supplier facing three buyers has less leverage than the scarcity argument alone suggests, which is why the litigation track exists alongside the commercial one.
The argument from input cost
Kopit Levien then attached a cost argument to the price test, and this is where the $2 billion figure does its work.
The company spent close to that sum in the prior year producing half a million works, she said, describing stories, photographs and videos as expressive creative work requiring significant resource and human beings to make. Against that, she set the spending of the model developers, who are committing tens of billions of dollars, and in some cases hundreds of billions, to talent, to compute and to power.
Her conclusion followed directly: the same discipline should apply to inputs. Model developers should be paying sufficient value for what they describe as data and what is in fact high quality independent journalism and other kinds of creative work.
The rhetorical move is worth noticing because it inverts the standard framing. Publishers usually argue from harm, pointing at referral traffic losses and lost advertising revenue. That argument has a weakness: it invites the response that the harm is speculative, or attributable to other causes, or a consequence of the publisher's own product decisions. An argument from input cost avoids that entirely. It says nothing about harm. It says that a firm spending $100 billion on graphics processors and electricity, while treating the corpus those processors read as a free input, has an accounting asymmetry that will eventually need resolving.
Whether the argument persuades a court or a counterparty is a separate matter. It gives negotiators a number where previously they had adjectives.
Why the per-unit figure is the durable part
Headline totals age badly. A $2 billion annual figure describes one company in one year, and any counterparty can argue that the number bundles costs unrelated to the corpus being licensed: printing, television production, games development, the commercial organisation.
The per-unit derivation survives that objection better. Close to $2 billion across roughly half a million works implies something in the region of $4,000 per work. That number is arguable in both directions, and the argument is productive rather than rhetorical. A licensee can reasonably contend that a wire-derived market report and a six-month investigation should not carry the same imputed cost. A licensor can reasonably contend that the investigation is only possible because the market report pays for the building.
Once the dispute is about the composition of a per-unit figure, it has become a commercial negotiation with a defined shape rather than a disagreement about principle. That is a meaningful change from where publisher and model developer discussions have sat since 2023, and it is the reason the disclosure carries weight beyond the company that made it. Any publisher can now perform the same calculation against its own cost base and enter a negotiation with a comparable, even if the resulting number is an order of magnitude smaller.
The figure also does work in the opposite direction, which is presumably understood. A publisher that cannot demonstrate a substantial per-unit production cost has just been given a benchmark it will struggle to meet, and the licensing market that develops around these reference points will not be equally hospitable to everybody in it.
Two tracks, run simultaneously
The publisher is pursuing litigation and commercial talks at the same time, and Kopit Levien described the split without apology. The company has sued three companies and has also done a deal.
The suits are against OpenAI and Microsoft, filed roughly two and a half to three years ago, and against Perplexity, filed in December 2025. The licensing agreement is with Amazon.
There are multiple ways to address the problem, she said, and enforcing rights in court is one of them. The stated objective is standing up to the proposition that these companies took the publisher's work and built products capable of competing with it.
The discovery record in the OpenAI matter has already produced consequences reaching beyond the two parties. A federal magistrate judge ordered OpenAI to hand over 20 million anonymised ChatGPT conversation logs to news plaintiffs in November 2025, and a stay request was denied on November 13 of that year. That production sets a precedent about what plaintiffs can extract from a model developer's operational records, and it lands well outside the newspaper business.
The Perplexity action sits inside a wider cluster. Encyclopaedia Britannica and Merriam-Webster sued the company in September 2025. CNN filed a complaint in May 2026 alleging the copying of more than 17,000 works, including verbatim reproduction of paywalled text through the Comet browser assistant. Ziff Davis sued OpenAI in April 2025 over content drawn from more than 45 properties. In September 2025, Anthropic agreed to a $1.5 billion settlement in a separate authors' case, the largest publicly reported figure attached to a copyright claim in the sector.
On the Amazon agreement, disclosure has been minimal. Chief financial officer Will Bardeen told analysts that affiliate, licensing and other revenues combine licensing deals, affiliate income, books, television, film and commercial printing, and can move unevenly between quarters, a point made on the August 5 earnings call. That combined line grew 7.1 percent year over year, which tells an analyst almost nothing about the licensing component and is presumably the intention.
What the tools actually do inside the newsroom
Kopit Levien separated compensation from internal deployment, and was considerably more specific about the second.
The company runs an AI initiatives team inside the newsroom. When the Epstein files were released on a Friday night, running to roughly three million pages, that team built a toolset to search the trove for patterns and clusters against reporters' standing questions. The result, in her description, was reaching information much faster and telling different kinds of stories than would otherwise have been possible.
Three further applications were described. During the transition into the second Trump administration, the tools reviewed the full public record of cabinet appointees more efficiently than manual reading allowed. In coverage of the Sydney Sweeney jeans advertising controversy, the newsroom used the technology to test the prevailing social media account of who was reacting, concluding that the reaction was first constructed by the political right. To document a shift in participation sports, existing aerial imagery of tennis courts was sorted at scale to count conversions to pickleball courts.
Each of those is a retrieval or classification task performed against material the company already holds or that is already public. None of them generates published prose. That distinction is the operating principle, and Kopit Levien stated it as a question about which work should be human led, answering that high quality independent journalism and creative work belong in that category. Reporting, in her account, is the foundation: a professional process of going into the world and unearthing new information, which then requires translation with judgment and care.
Automated voice and translation supply the accessibility applications. Coding tools have reached a large share of product development and marketing staff. On token spend as a planning line, she said every chief executive she knows is discussing it and nobody has a credible answer yet, a remark that connects directly to a separate thread running through the same week.
Roughly a thousand people work in digital product development at the company and about three thousand in journalism and content production. Her expectation, conditional on continued commercial success, is that the second group keeps growing.
Headcount, coverage cost and the advertising claim
The newsroom holds the largest collection of journalists in the company's history, about a thousand more than when Kopit Levien joined thirteen years ago, excluding The Athletic and the Wirecutter newsroom.
Departures were raised directly, including columnist Ross Douthat's move to 60 Minutes and the exit of the Hard Fork co-hosts. Her answer was structural rather than financial: the model rests on editors, lawyers, security staff, graphics and photography, and an audience that changes what publication means, not on individual compensation.
Cost of coverage was used to describe how the commercial side serves the editorial side. About 70 people were deployed in and out of Ukraine last year, five years into the war. Sending a correspondent into an Ebola zone is expensive. Her stated job is keeping the company successful enough that no editorial leader has to weigh cost against public interest.
That position was extended to advertising without qualification. On Wirecutter, the product review operation, she said the recommendation follows the testing even when the losing mattress company is the large advertiser. Every premium publisher selling against editorial adjacency has to be able to make that claim. Not all of them make it flatly.
Video, and the difference from 2015
Weisenthal put the earlier industry pivot to video directly to her, noting that the company's stock sold off after each of the last two earnings reports and that video production carries cost.
Kopit Levien ran the commercial business during the first pivot and characterised it as a cynical attempt to get advertising dollars, made without regard to whether millions of people would actually engage. That is an unusually blunt retrospective assessment of an industry-wide episode that cost a large number of journalists their jobs.
Her case for the current investment rests on audience rather than inventory. Reading and listening audiences continue to grow, and video reaches a separate population that watches news or news-adjacent entertainment on YouTube, TikTok and Instagram. The push is being staffed largely from people already at the company, with capability added later, and most podcasts have been converted into video shows.
The advertising consequence is supply. The publisher attributed part of its second quarter digital advertising growth to growth in advertising supply, and video formats are one mechanism that creates it. For a media buyer, that is the operative sentence in the whole interview: a premium publisher expanding sellable inventory in a quarter when several peers reported the opposite.
The demand curve, and what it protects
On subscription pricing, Kopit Levien said the company never accepted a choice between wide availability for sampling and enough value in the paid product to convert. The stated strategy runs from discounted international games subscriptions at one end to shared family bundles at the other, designed to get everybody under the demand curve.
The entry mechanic is a bundle at a dollar a week, followed by engagement, followed by a step-up to a materially higher price once value has been experienced. Digital-only average revenue per user reached $9.94 in the second quarter, up 3.1 percent year over year.
She also described a decade-long orientation away from platform distribution, saying the company recognised for the better part of a decade that the direction of travel from the platforms meant more of the relationship had to move into its own control. That decision, made years before AI answer engines existed, is the reason the company is negotiating from a position that most publishers no longer occupy.
The contrast in the same reporting window is stark. The publisher reported second quarter digital advertising revenue of $114.0 million, up 20.7 percent, with total revenues of $762.5 million and 13.35 million subscribers. Peers moved the other way. USA TODAY Co. recorded digital advertising down 9.2 percent and lost 22 million monthly unique visitors in a single quarter. Ziff Davis took a $54.8 million goodwill impairment against its health division as advertising revenue fell 6 percent. News Corp reported advertising falling to 15 percent of total revenue across a $9.03 billion year.
The backdrop against which all of those numbers sit is referral decline. Chartbeat data showed small publishers lost 60 percent of search referral traffic over two years. The first randomised study of AI Overviews measured a 39.8 percent reduction in outbound clicks. One independent site reported a 75 percent traffic loss as zero-click search rates climbed.
The asset underneath both negotiations
For media buyers, the addressable consequence of all of this is first-party data.
The publisher's first data clean room collaboration, disclosed on June 21, 2026, reported a 61 percent increase in international click-through rate and a 25 percent performance uplift in the United States against cookie benchmarks. A logged-in base of 13.35 million is the asset that the licensing negotiation and the advertising sales operation both rest on, which is why the three conditions Kopit Levien described are not only about AI. They describe the terms on which a proprietary audience and a proprietary corpus are rented out to anybody.
The sports operation has been extended along the same lines. The Athletic gained its first connected television home in July 2026, when six of its shows were opened to sponsors on Fubo. The Athletic was acquired with roughly 450 journalists and a small business team, many recruited from local newspapers. The newsroom is now about a hundred people larger, Steven Ginsburg was hired from The Washington Post about four years ago, and a substantial editing layer was added.
On further acquisitions, Kopit Levien did not rule anything out but set a limit, noting that the bar for return on investment is high and that Wordle cost considerably less than The Athletic, with the rest of the games portfolio built internally.
The crawlers that will not identify themselves
If the licensing question is about price, the enforcement question is about identity. Both surfaced in the same 48 hours.
On August 10, 2026, AdExchanger reported that three members of the United States House of Representatives had introduced the Stealth Bot Prohibition Act, a bill that would require AI stealth crawlers to disclose their identity and purpose to the host of a website. The sponsors are Valerie Foushee, Democrat of North Carolina, and Laurel Lee and Gus Bilirakis, both Republicans of Florida. The legislation was authored by the News/Media Alliance, a trade association representing over 2,000 media brands, and introduced in late July.
A version of the act passed in New York state earlier in 2026, which is the detail that makes the federal bill more than symbolic. State-level precedent gives the sponsors a working text and a demonstration that the drafting survives contact with a legislature.
The mechanics, and the number attached
The bill would allow enforcement agencies including the Federal Trade Commission and state attorneys general to bring cases involving undisclosed crawlers to court. Each violation would incur a fine of $53,000.
That per-violation structure is the operative design choice. A crawler that fetches a million pages while concealing its identity is not committing one violation under a per-request reading, and the arithmetic becomes existential quickly. Danielle Coffey, president and chief executive of the News/Media Alliance, framed the hope as steep penalties combined with active monitoring by website owners and enforcers deterring malicious scraping.
Coffey also described the underlying commercial problem in terms that will be familiar to anyone who has watched an arbitrage market form. These bots are creating a reseller market by crawling and scraping publisher sites, then serving or selling the content elsewhere without permission.
Amelia Binder, senior vice president of global government affairs at Axel Springer, put the identity question plainly: a bot trying to hide what it is probably is not operating above board. Most stealth bots on Axel Springer properties have been scraping content against terms of service for financial gain, either passing it off as their own or reselling it to third parties, which produces both revenue loss for the publisher and higher odds of inaccuracy as content gets regurgitated out of context.
"We need this first line of defense," Binder said, describing the bill as a first important step toward transparency and licensing deals.
Why robots.txt stopped working
The technical explanation for why legislation is being contemplated at all is worth stating precisely, because it is a story about a voluntary standard reaching the end of its useful life.
Web crawling has historically run on an honour system administered through public robots.txt files. Publishers host those files. Crawlers read them. The files indicate which parts of a site may be accessed. There is no enforcement mechanism of any kind; the files are recommended guidance that well-behaved crawlers choose to follow.
Recently, crawlers have begun ignoring and in some cases manipulating those files, masquerading as human viewers so that sites categorise the requests as human traffic rather than bot traffic. Once inside, the operator ingests the content and either passes it off as its own, resells it to third parties, or supplies it as training data for models and the APIs built on top of them.
Lark-Marie Anton, chief communications and brand officer at USA TODAY Co., attributed the regulatory gap to speed: AI development has outpaced existing regulation. The gap is real but the framing understates the problem. Robots.txt was never regulation. It was a convention that worked while every meaningful crawler operator had a commercial reason to be identifiable, because the crawler was attached to a search engine that sent traffic back. Remove the traffic exchange and the convention loses its incentive structure. Nothing was broken. The reciprocity underneath it simply stopped existing.
Several dozen companies, including Perplexity, have been profiting from sales derived from illicitly scraped data, per reporting the article cites from Digiday.
The costs publishers are already carrying
The financial detail in the reporting is more concrete than the legislative outlook.
Twenty-five percent of Politico's hosting costs now go toward bot management, Binder said, an expense the company did not budget for. Axel Springer owns Politico and Business Insider among numerous other publications. A quarter of infrastructure spend redirected to managing traffic that generates no revenue is a material margin event for a business already absorbing referral decline.
Conan Gallaty, chief executive and chairman of the Tampa Bay Times, described a service-quality consequence: unauthorised bot activity weighing down networks such that human visitors have trouble accessing articles. The phrasing was that it taxes the publisher's ability to serve customers.
This is the same phenomenon documented at much smaller scale elsewhere. One operator blocked Amazon's AI crawler after recording 117,000 daily page reads, with Anthropic's crawler hitting a 35,000 to 1 crawl ratio and CAPTCHA solve rates measuring 0.24 percent. The economics scale down without improving: a small operator absorbing infrastructure cost for machine reads that produce no advertising impression, no subscription and no referral.
The reseller market, and why it prices at zero
The phrase Coffey used deserves more weight than it usually receives. A reseller market is not a metaphor for theft. It is a specific market structure, and it has specific economics.
In a functioning content market, the producer sets a price and the distributor takes a margin. In a reseller market built on undisclosed scraping, the input cost to the reseller is bandwidth, and nothing else. The reseller does not commission reporting, does not carry legal risk for what the reporting alleges, does not employ editors, and does not pay for the correspondent flown into a conflict zone. It resells an artefact whose production cost sits entirely on somebody else's income statement.
That asymmetry sets the clearing price of the resold artefact near zero, because any reseller can be undercut by another reseller with the same input cost. The consequence is not that publishers earn less from the reseller. It is that the resold version competes with the original at a price the original can never match, while the original absorbs the infrastructure cost of supplying it.
Quantifying that cost has been the missing step. The IAB Tech Lab has been working the problem directly, asking content owners to establish how much bot traffic is actually costing them. The Politico figure gives the exercise a benchmark: a quarter of hosting spend. Applied across a publisher estate of any size, that is a number that belongs in a licensing negotiation, because it converts an abstract complaint into a line item that the counterparty's own crawler is generating.
There is a related development on the monetisation side that reframes the same traffic as an opportunity rather than a cost. AdExchanger has documented publishers finding new ways to monetise their data through AI agents, which depends entirely on the agent being identifiable. An anonymous crawler cannot be charged, cannot be rate-limited by contract and cannot be offered a premium tier. Disclosure is therefore not only an enforcement mechanism. It is the precondition for any pricing mechanism at all, which is the strongest commercial argument for the bill and the one least likely to appear in the floor debate.
The security argument and the political weather
Coffey raised a national security dimension that explains the bipartisan sponsorship. Many of the bots originate from Russia and China, and harvesting news and information at scale allows the operators to sift it for patterns, train their own systems and undercut the business models supporting independent journalism.
Gallaty added a jurisdictional point: intellectual property needs to be respected outside national borders, and when these bots scrape content they do whatever they wish with it.
Whether that argument carries the bill is uncertain. The current administration has generally opposed stricter AI regulation on competitiveness grounds, though it is finalising an AI oversight framework amid increasing concern about electricity costs and cyberattack potential. Gallaty reported interest from both sides of the aisle, and Coffey placed her hope in the bill's bipartisan appeal and the fact that it is reasonable and on its face unobjectionable.
There is also a second-order effect that cuts the publishers' way. The flood of AI-generated content has increased demand for trusted, high-quality journalism, Anton said, a dynamic that could lead to licensing and attribution frameworks that actually pay publishers. That is the same commercial logic Kopit Levien priced at close to $2 billion a year, arriving from the opposite direction.
Read against each other, the two stories describe a two-track strategy that the larger publishers are running deliberately. Litigation and legislation establish that permission is required. Licensing establishes what permission costs. Neither track works alone. A price with no enforcement is a suggestion, and enforcement with no price is a blockade.
Impressions arrive. Clicks do not
The measurement layer for machine-read content moved on August 11, 2026, and moved incompletely.
Search Engine Roundtable reported that the Google Search Console generative AI performance report appears to be live for everyone, based on a check across all profiles, with no formal announcement of full rollout from Google. Access had been expanded to a wider group a few weeks earlier. Vijay Chauhan flagged the broader availability on X that morning.
The report sits as an expandable tab under the main performance report. It carries five dimensions.
Impressions record how often URLs from a site appeared in generative AI features across Search and Discover. Pages identify which specific URLs appeared. Countries break visibility down by market. Devices identify what people were using when they saw the site, available for Search results. Dates support hourly, daily, weekly and monthly granularity.
The absences are the story. There is no click data and no query data.
What can and cannot be built on this
A site owner can now answer several questions that were unanswerable in July. Whether generative impressions are rising or falling over time. Which pages earn the most and fewest generative impressions. Which countries and devices generate that visibility.
A site owner still cannot answer the two questions that determine whether any of it is worth funding. What did people ask that surfaced the page, and did anybody click through.
For anyone attempting to price content production against machine consumption, that gap is the whole problem restated as a product limitation. An impression inside a generated answer is a use of the work. It is now countable. It remains unmonetised and, in the absence of query data, unoptimisable in the way search impressions have been optimisable for two decades. The report tells a publisher that the machine read the page. It does not say what the machine was asked or whether the reader ever arrived.
The measurement asymmetry mirrors the commercial one. Model developers know exactly what they spent on compute. Publishers now know approximately how often they were read. Neither figure connects to a payment.
The precedent for counting without connecting
There is a useful historical parallel for a metric that counts exposure without connecting it to outcome, and it is not a comforting one.
For most of the past two decades, search marketing has operated on a closed loop. A query produced an impression, an impression produced a click, a click produced a session, and a session produced a measurable outcome. Every optimisation practice in the discipline assumes that loop. Keyword research assumes query data. Landing page work assumes arrival. Bid strategy assumes a downstream conversion signal to bid toward.
The generative report breaks the loop at both ends simultaneously. Without query data, there is no input to optimise against. Without click data, there is no outcome to optimise toward. What remains is a volume count of an event the site owner did not initiate, cannot influence through any established technique, and cannot convert.
That is closer to broadcast reach measurement than to search analytics, and the comparison is instructive rather than reassuring. Broadcast reach was monetisable because a currency existed, agreed between buyers and sellers, and because the seller controlled the inventory. In the generative case, the party being measured does not sell the placement, does not set its frequency, and receives no payment when it occurs. A currency without a transaction is a statistic.
Third-party vendors have moved into the gap. Cloudflare launched scoring of brand visibility inside Claude and GPT answers in early access, with four metrics covering citation, mention and prominence rates for site owners, though whether probing only two assistants measures enough of the field remains an open question. The commercial logic behind that product launch and behind the Search Console report is identical: visibility inside generated answers is becoming a category that brands will pay to measure, well before anybody has established that it is a category brands can influence.
For marketers, the practical distinction is between a diagnostic and a control. The generative report is a diagnostic. It will show a publisher or a brand whether its presence inside machine answers is growing or shrinking, which is genuinely new information. It will not show why, and it will not show what that presence is worth. Both of those remain unpriced, which is the same condition the licensing negotiation, the crawler bill and the token budget are each trying to resolve from a different angle.
The crawling map keeps expanding
The same reporting window carried a related signal about who is building an index at all.
Search Engine Roundtable reported on August 10, 2026 that Meta is crawling the web and may be building its own search engine, based on screenshots posted by developer Pieter Levels on August 6 showing crawling activity across his properties. Meta operates Facebook, Instagram, WhatsApp, Messenger and Threads. The company has wanted its own search service for well over a decade, partnered with Bing to power web search features, and dropped that partnership a couple of years later.
A new large-scale crawler entering the field changes the arithmetic for every publisher weighing infrastructure cost against machine traffic. It also changes the negotiating position of anybody currently licensing to a small number of model developers, in both directions. More potential counterparties means more competition for a corpus. More crawlers means more hosting cost carried before any of them pay.
Separately, Search Engine Roundtable reported that Cloudflare is attempting to give sites a way to block Google AI Overviews. The technical difficulty there is well understood: the crawler that indexes a page for search results and the system that composes an AI answer have historically shared infrastructure, which means blocking one has tended to mean blocking both. That coupling is the reason most publishers have not exercised the option they nominally hold.
Google also expanded AI labelling into new ad surfaces. Search Engine Roundtable documented AI labels appearing on Local Pack and Google Discover ads on August 10, extending disclosure into placements where the label had not previously appeared.
The threshold doubles
The creator economy received its own repricing on August 10, 2026, and the direction was upward.
MediaPost reported that YouTube doubled the eligibility requirements for creator monetization. Creators signing up for the YouTube Partner Program will need to show at least 8,000 qualified watch hours over the past 365 days, or 20 million qualified Shorts views over the past 90 days, before they can begin monetizing.
The prior thresholds were 1,000 subscribers and 4,000 watch hours over the past year, or 1,000 subscribers and 10 million Shorts views over the past 90 days. The watch-hour requirement and the Shorts view requirement have each doubled. The change takes effect on February 1, 2027.
These are the first major changes to the revenue-sharing program since 2018, when eligibility moved from 10,000 views per channel to the model now being replaced. The YouTube Partner Program is 20 years old and holds over 3 million creators.
The justification and the arithmetic
YouTube's stated rationale is keeping pace with the growth of the platform, which now records over 200 billion daily Shorts views and over a billion hours of watch time on television per day.
The logic is internally consistent. If aggregate consumption has multiplied, a fixed threshold admits an ever-larger cohort to revenue sharing, and the average revenue per participating channel falls. Doubling the bar restores something closer to the original selectivity.
The distributional consequence runs the other way. Doubling the view requirement makes it harder for creators to earn money from advertisements, and disincentivises creators with niche followings in particular. That subset matters commercially well beyond the creators themselves: smaller brands frequently use niche creators as the entry point into specific communities, precisely because those creators are affordable and their audiences are coherent.
A creator who would have qualified at 5,000 watch hours in January 2027 does not qualify in February. Nothing about that creator's audience changed. The floor moved.
What the doubling does to the middle of the market
The arithmetic is worth working through, because the headline framing of a doubled threshold understates the effect on the Shorts path specifically.
On the long-form path, a creator previously needed 1,000 subscribers and 4,000 watch hours. The new requirement is 8,000 watch hours. Four thousand hours across a year is roughly eleven hours of watch time a day, achievable for a channel with a small but committed audience. Eight thousand hours is twenty-two hours a day, every day, for a year. That is not a marginal adjustment to a hobbyist channel. It is a different class of operation.
On the Shorts path, the move from 10 million to 20 million qualified views over 90 days is steeper still in practice, because Shorts views are cheap individually and volatile in distribution. A creator whose 90-day window happens to contain one algorithmically favoured video can clear a threshold that a creator with steadier output cannot. Raising the bar amplifies that variance rather than smoothing it, because the required total moves further from what consistent posting produces and closer to what a single distribution event produces.
The February 1, 2027 effective date gives existing applicants roughly six months, which is a reasonable notice period and also, for a channel currently sitting between the old and new thresholds, an uncomfortable one. Six months is long enough to try and short enough to fail.
For advertisers, the consequence is compositional rather than immediate. The pool of monetising channels will not shrink, since existing partners are not being removed. It will stop growing at the bottom. Over several years that produces a partner base skewed further toward scale, which suits brands buying reach and works against brands buying specificity. Smaller advertisers that rely on niche creators as an affordable entry into defined communities will find the supply of newly monetising niche channels thinner than the platform's overall growth would suggest.
Where this connects
The threshold change lands alongside an expansion of what a qualified creator can earn from. YouTube opened its Shopping affiliate program to UK creators with Wayfair, Currys, Debenhams, Boots, Marks and Spencer and Etsy as launch merchants, making the United Kingdom the fifteenth market, with commissions clearing 60 to 120 days after purchase.
The combined effect is a narrower gate opening onto a larger room. Fewer creators cross the line; those who do have more revenue mechanisms available once across. That is a deliberate concentration of the creator base, and it points the same direction as everything else in this edition. Value is being consolidated toward whoever already holds scale, whether the unit is a newsroom, a corpus, an audience or an inventory pool.
Google also extended measurement to creators without websites. A July 29 guide brought Search Console query, country and device breakdowns to TikTok, Instagram, X and YouTube data, while flagging what a YouTube playlist filter does not measure.
When a category becomes plumbing
The most quietly consequential item of the three days involved a brand disappearing rather than a number changing.
Digiday reported on August 11, 2026 that Experian has retired the Audigent brand, folding it into Experian Marketing Services, twenty months after acquiring the company at the end of 2024. An Experian spokesperson framed the change as consolidation rather than retreat, saying Audigent's curation capabilities remain an important part of the platform and a critical component of the broader technology and growth strategy, and that the vision of embedding intelligence into every supply-side decision has not changed. Whether the change leads to layoffs, Audigent declined to say.
What curation used to mean
The reason this matters extends past one brand.
When the acquisition was done, curation had an agreed definition. Specialists sat between demand-side platforms and supply-side platforms, packaging audience signals, contextual signals and supply-path signals into deal IDs that made overlooked inventory sellable. The service was discrete. It could be named, bought, and invoiced.
That legibility is gone. Curation is no longer a category with an agreed definition. It is a label that everybody in the supply chain has claimed and bent toward whatever they happen to sell.
Supply-side platforms sell curation as something they power rather than something they perform. Demand-side platforms pitch it as evidence that they do not need the sell side doing their targeting. Identity and data companies, Experian included, frame it as a component of a broader graph rather than a discrete product. Holding companies folded it into supply-path optimisation. Ask five firms what curation means and the answer depends entirely on what each of them is selling.
Marc Fanelli, senior vice president of global digital audiences and operations and chief operating officer at Eyeota, offered the historical parallel. Data management platforms and identity followed the same arc. Those categories did not disappear because they stopped being valuable; they disappeared because their capabilities became embedded throughout the ecosystem.
"It's becoming infrastructure," Fanelli said of curation.
The same pattern, priced differently
This is the pricing problem again, arriving from a different direction.
A discrete category can be sold at a margin. Infrastructure is a cost centre that competitors expect to be included. When curation was a category, a curator could charge a fee for run-time evaluation of audience, content and supply-chain signals. When curation is infrastructure, that same evaluation is a feature that a supply-side platform includes to win the integration, and the fee compresses toward zero.
The Digiday reporting on advertiser behaviour in the same window shows what that looks like from the buy side. Georgia-Pacific cut its supply-side platforms by 80 percent, having set out to stop paying roughly thirty intermediaries that all appeared to be selling more or less the same inventory. The advertiser expects further gains as the algorithms powering its curation improve, and has been explicit that the consolidation will not come at the expense of its demand-side platform, Yahoo. For that to work at scale, advertisers need supply-side platforms to be considerably more transparent about how they sell, who they use to do it, and what they take in fees along the way.
The sell side is moving in the mirror direction. On Magnite's second quarter earnings call, chief executive Michael Barrett described a company adding demand-side capabilities for planning and activating campaigns without becoming a demand-side platform. SpringServe is being used not only to serve ads but to package inventory and audiences and to handle optimisation decisions, all of which historically happened on the buy side. Total second quarter revenue was $193 million, up 11 percent, with contribution excluding traffic acquisition costs at $190 million, up 17 percent, and connected television contributing $97 million, up 36 percent.
Barrett's framing of the transition was direct. "But agents do not eliminate infrastructure; they increase the need for it," he said, arguing that as publishers and advertisers lean harder into data and AI, audience enablement and decisioning will inevitably move from the buy side to the sell side.
Commerce media illustrates the mechanism. Fanatics, CVS Media Exchange, Best Buy and PayPal Ads use Magnite to plug first-party data into connected television and open internet campaigns. Walmart Connect combines Walmart commerce data with Vizio inventory to support off-site campaigns and provide closed-loop measurement. The stated principle is that no single demand-side platform should get privileged access to data and inventory.
Barrett also delivered the sentence that connects this section to the rest of the edition, observing that the days of making an easy living by stringing demand-side platforms together and selling banners are ending, and adding that this bodes well for Magnite though perhaps not for the whole ecosystem.
Curation dissolving into infrastructure and intermediaries being cut by 80 percent are the same event observed from two positions. A layer that used to charge for a service is becoming a layer that has to justify its existence against the alternative of being absorbed.
The protocol layer arriving underneath
There is a second reason the curation brand stopped being necessary, and it has less to do with definitional drift than with standardisation.
When a capability is bespoke, it needs a vendor. When it is specified, it needs an implementation. Agentic transaction standards have been moving from proposal to specification through 2026, and Digiday's explanation of the Ad Context Protocol as the blueprint for how AI agents transact describes the direction: a common vocabulary through which buying and selling systems can negotiate without a human translating between them. A standard that specifies how audience, context and supply-path signals are expressed is, functionally, a standard that specifies what a curator used to be paid to assemble.
The tooling is following the specifications rather than the other way round. Prebid.js merged a DevTools module allowing AI agents to read live header bidding auctions from Chrome after 19 commits and 109 passing checks, though it remains experimental and behind a flag. Publishers being able to query auction state programmatically, and agents being able to do the same, removes another category of intermediary knowledge that used to be sold as a service.
Magnite launched Magnite Orchestration in June as a software layer helping agents communicate and coordinate, and has also described two routes for partners to run their own AI models inside its auction, with Chalice AI and inPowered AI embedding decisioning engines inside the supply-side platform alongside Omnicom Media and SWYM.ai. That is the infrastructure position stated as a product: the exchange does not sell the decision, it sells the place where somebody else's decision executes.
Fee compression follows mechanically. A curator charging for assembly competes against a specification that describes assembly for free and an exchange that hosts execution as a cost of doing business. The margin does not migrate to a different intermediary. It disappears into the integration, which is what Fanelli meant, and what the retirement of a well-known brand into a parent company's marketing services division looks like from the inside.
None of this settles who captures the value released by that compression. Georgia-Pacific's 80 percent reduction in supply-side platforms suggests the advertiser captures some of it. Magnite's 17 percent growth in contribution excluding traffic acquisition costs suggests the exchange captures some of it. The parties that captured none of it are the ones whose entire business was the layer that dissolved.
The meter under the work
If publishers are trying to price their inputs, agencies spent the same week trying to price theirs.
Digiday reported on August 11, 2026 that rising AI token costs are forcing marketing services groups to confront how much they should be spending on AI. Token usage is exploding at Monks, in the description of co-founder and chief AI officer Wesley ter Haar, a consequence of the company's technology client base, production heritage and years-long investment in generative tools.
The financial context is unusually favourable, which is what makes the cost question interesting rather than alarming. Parent company S4 Capital published first half earnings the previous week that prompted a share price jump for the London firm. Strict cost discipline broadened margins to 12.3 percent and doubled first half operating profit to £35.2 million, up from £16.4 million in the same period a year earlier. Those figures followed a positive set of first quarter numbers.
S4 chair Sir Martin Sorrell called the result a huge improvement while acknowledging that rising token costs could demand a stricter spending policy.
The shape of the problem
Sorrell cited a Forrester survey finding that 60 percent of agencies were prioritising spending on third-party tools, generally paid for through a seat or software-as-a-service model, against 35 percent prioritising cloud or compute costs.
That split is the crux. A seat licence is a fixed, forecastable cost that behaves like headcount. Compute is a variable cost that scales with usage and cannot be forecast reliably, because the usage depends on how enthusiastically staff adopt tools that the business has just spent two years encouraging them to adopt.
Sorrell's assessment of his own group was candid. "It's probable that we've overused or overindulged," he said, while adding that the reality is the company will have to spend more on increasing capability.
Ter Haar proposed a differentiated approach, with caps set for teams using agents or tools for lower priority tasks and lifted for areas of genuine staff expertise. That is a rationing design rather than a pricing design, and it concedes the underlying difficulty: nobody has yet established what an hour of machine-assisted work is worth to a client, so the industry is managing the input cost while the output price remains unsettled.
The parallel with the publisher side is exact and probably not coincidental. Kopit Levien said every chief executive she knows is discussing token spend as a planning line and nobody has a credible answer yet. Sorrell, running a very different business, arrived at the same position in the same week. One side is trying to establish what its human-produced inputs are worth to machines. The other is trying to establish what its machine-produced outputs are worth to clients. Neither has a settled number.
Three answers to the same unsettled question
S4 is not alone in the problem, and the industry has produced at least three incompatible responses to it.
Digiday's earlier survey of how ad agencies are grappling with AI costs as spending outpaces demonstrated value laid out the split. Dept declines to pass token costs to clients at all, arguing that itemising them cheapens the value of the person using the tool by making the spent tokens the metric rather than the work produced. Monks has gone the other way, building tokens directly into its technology and subscription pricing. The large holding companies have largely tried to avoid the question, folding AI costs into broader commercial structures such as principal media arrangements.
Rationing has emerged as the practical middle path. PMG rolled out a company-wide tool pooling staff access to major models under a $50-a-day token cap per user, following months of alpha and beta testing that began in January when testers ran without limits. Before the cap existed, staff queried models directly and unmonitored, and the costs accumulated.
All three approaches stand or fall on the same missing element, which is a way to define and measure what the technology actually delivered. The evidence on that point is not encouraging. KPMG found that 49 percent of organisations cut AI agent rollouts when costs outran value across 2,145 leaders surveyed, with established return on investment sitting at 7 percent while planned AI spending held at $188 million. Cost visibility, on that reading, is what separates the firms seeing returns from the firms discovering they were not.
The compensation model is moving in the same direction and at the same pace, which is to say slowly. WPP chief executive Cindy Rose told analysts that outcome-based pay remains years away, with her stated priority being a return to positive organic growth. Outcome-based pricing is the structure that would resolve the token problem cleanly, since a client paying for a result is indifferent to the compute consumed producing it. Everybody in the industry agrees it is the endgame. Nobody has moved there, because it requires an agreed definition of the outcome, which is the same measurement gap that leaves the generative impressions report without click data and leaves a licensing negotiation without an agreed unit.
The exit from public pricing
The final thread is about where prices get set at all.
Digiday's Ad Tech Briefing on August 11, 2026 examined the growing expectation that more listed ad tech companies will go private, noting that public markets have become an uncomfortable place for many companies in the sector and that the latest quarterly earnings underlined the point. The largest public names are plateauing, which is prompting further speculation.
The pattern already has precedent. Integral Ad Science and LiveRamp both left or are leaving the public markets, and LiveRamp faces an August 17 shareholder vote on the Publicis buyout at $38.50 a share, with closing still set for before the end of 2026 and revenue at $214 million in the most recent quarter.
What changes when these companies go private is disclosure. A listed measurement or identity vendor files quarterly, discloses revenue concentration, describes pricing pressure in risk factors and takes questions from analysts on a recorded call. A private one does none of that. The industry loses its clearest external read on what the middle of the supply chain is actually earning.
That matters for the pricing thread running through this edition because the middle of the supply chain is where the disputed value sits. Advertisers arguing about supply-path fees, publishers arguing about verification costs, and agencies arguing about technology charges have all historically used listed-company filings as reference points. Remove enough of those filings and the arguments become assertion against assertion.
The same window carried a related repricing on the demand side. Digiday reported that The Trade Desk has restructured Identity Alliance payouts around incrementality, shifting the basis on which identity partners are compensated. Changing the payout basis from volume to measured contribution is, in effect, a decision that a previously assumed value now has to be demonstrated. The same instinct is visible everywhere in this edition.
What disappears with the filings
It is worth being specific about which disclosures matter, because not all of them do.
Revenue concentration disclosure matters. A listed vendor has to say when a single customer represents a material share of revenue, which tells the rest of the market how much pricing power that customer holds. Risk factor language matters, because it is the one place a company is legally obliged to describe pressures it would otherwise present as opportunities. Segment reporting matters, because it reveals which parts of a bundled offering are actually carrying the business. Analyst calls matter least in content and most in tone, since the questions asked reveal what institutional holders have concluded.
Take a vendor private and all four vanish at once. What replaces them is a marketing narrative, and the counterparties negotiating fees with that vendor lose their only independent reference for whether the quoted rate is close to the rate everybody else pays.
The timing compounds the effect. This is happening precisely as advertisers push for more granular fee transparency. Georgia-Pacific's requirement that supply-side platforms be considerably more open about how they sell, who they use and what they take along the way is the direction of travel on the buy side. Public filings were never a complete answer to that demand, but they were an external check that did not depend on the vendor's willingness to answer.
There is a second-order consequence for the pricing thread running through this edition. Every negotiation described here works better with a public reference point. The licensing negotiation improved the moment a publisher put a production cost on the record. The bot enforcement argument improved the moment a publisher disclosed that a quarter of hosting spend goes to bot management. The token debate is being conducted with published margin figures and a named daily cap. The middle of the ad tech supply chain is moving in the opposite direction, toward fewer public numbers at the point where the disputed money sits.
What the three days actually establish
Across August 10 to August 12, 2026, six developments arrived from six directions and converged on one unresolved question.
A publisher priced its annual output at close to $2 billion and asked model developers to reason about inputs the way they reason about compute. Three legislators proposed making anonymous crawling actionable at $53,000 per violation, with a state precedent already on the books. A search platform gave publishers a count of machine reads without the click data that would make the count monetisable. A video platform doubled the volume threshold at which a creator starts earning. A data company retired the brand that had defined an entire ad tech category, because the category had become infrastructure nobody bills for separately. An agency group with doubled operating profit conceded it had probably overused the technology it sells.
Each of those developments was reported as a discrete item by a different outlet, and each was framed in the vocabulary of its own sector. The publisher story was framed as copyright. The bill was framed as regulation. The Search Console change was framed as a product update. The YouTube change was framed as creator economics. The Audigent retirement was framed as consolidation. The token story was framed as agency margin.
Read in sequence across three days, the sector framings fall away and the same structure appears underneath each one. In every case, a party that produces something is discovering that the party consuming it has no established obligation to pay, no agreed method to count, or no external reference against which to argue about the number. The publisher produces journalism consumed by models. The website operator produces pages consumed by crawlers. The creator produces video consumed by a platform that sets the eligibility bar. The curator produced assembly consumed by a specification. The agency produces machine work consumed by clients who have not agreed what it is worth.
The connective tissue is not AI as such. It is that the industry has spent two decades building measurement and pricing systems for human attention, and is now transacting in machine attention without an agreed unit, an agreed rate or, in several cases, an agreed way to count it.
What has changed in the past three days is the supply of usable numbers rather than the resolution of any dispute. Close to $2 billion against half a million works. Fifty-three thousand dollars per undisclosed crawl. Twenty-five percent of one publisher's hosting spend. Eight thousand watch hours from four thousand. Twelve point three percent margin against a doubled operating profit. Eighty percent fewer supply-side platforms at one advertiser. Fifty dollars a day per user at one agency.
None of those figures settles anything on its own. Collectively they represent more disclosed arithmetic about the economics of machine-consumed content than the trade press produced in the preceding quarter, and disclosed arithmetic is what negotiating positions are built from. The parties that put numbers on the record this week did so because they had concluded that an unpriced input eventually becomes a free one.
Kopit Levien's three tests are a reasonable place to watch from, because they generalise past their original subject. Continuity, control, and a sustainable value exchange are the questions a publisher asks a model developer. They are also the questions an advertiser asks a supply-side platform, a creator asks a video platform, and a client asks an agency about its token bill. In every one of those relationships this week, at least one party discovered that a number they had been treating as settled was not.
Also noted
- August 11, 2026 - iHeartMedia takes six podcasts to Disney+ and Hulu: Six iHeartPodcasts titles reach the streaming platforms under a weekly arrangement beginning August 14, with Hey Jonas! on both services and Pod Meets World on Disney+.
- August 11, 2026 - Trump Media reports a $238 million quarterly loss: The company behind Truth Social attributed the loss to entry into crypto and online betting businesses unrelated to media, more than ten times the loss recorded in the same period a year earlier.
- August 10, 2026 - OpenAI targets smaller advertisers: Job listings indicate the company is building toward the small and medium business advertiser base that supplies the majority of ad revenue at the largest platforms, following the arrival of Google Shopping-style product carousels in ChatGPT ads.
- August 10, 2026 - Google Local bars repeated bilingual names: Business Profile guidelines now disallow repeated bilingual names and transliterations in listing titles, closing a keyword-stuffing route used in multilingual markets.
- August 10, 2026 - Google clarifies hreflang indexing behaviour: Alternate language URLs declared through hreflang are not indexed in the conventional sense, a distinction that affects how international sites read their coverage reports.
Discussion