The crawler that will not give its name
More than 300 publishing executives went to Washington last week to lobby for a single technical requirement: that an automated agent reading a website should say what it is and why it is there. Digiday reported the trip on September 29 at 04:01 UTC, naming the bill they came to support as the Stealth Bot Prohibition Act, introduced in July 2026. The legislation would oblige AI scraping bots to identify themselves and disclose their purpose instead of disguising their traffic as something else.
The delegation was assembled by the News/Media Alliance and America's Newspapers, and its roster was unusually senior. Roger Lynch, chief executive of Conde Nast, made the trip. So did Debi Chirichella, president of Hearst Magazines and board chair of the News/Media Alliance, who described the value of the exercise as coming together and speaking with one voice. Mike Reed, chief executive of USA Today Co., called the bill a meaningful step. Pat Dorsey, publisher of the Santa Fe New Mexican and chair of America's Newspapers, attended alongside Alan Fisco, president and chief executive of The Seattle Times and president of America's Newspapers. Danielle Coffey, president and chief executive of the News/Media Alliance, noted that this year's fly-in carried roughly double the executives of the 2025 edition and ran an extra day of programming.
Why now? The numbers that publishers brought with them have moved fast. Bots account for more than 60% of overall internet traffic. TollBit counted more than 22 billion AI bot scrapes across the sites it analysed in the first half of 2026. The ratio of AI bot visits to human visits rose sixfold inside a single year, from one in 200 in the first quarter of 2025 to one in 31 in the fourth. PPC Land has tracked the same measurement problem from the crawler side, reporting that 15% of AI page fetchers in Europe reached URLs their targets had disallowed, and that Cloudflare's chief executive expects bot traffic to reach a thousand times human traffic within five years.
The federal bill has a state precedent. New York passed its own Stealth Crawler Prohibition Act in June 2026, which obliges AI crawlers to identify themselves to news sites. The federal draft is broader: it reaches websites and digital platforms generally, not only news publishers. David Buttle, founder of the publisher coalition SPUR, suggested that copyright litigation would keep reinforcing the pressure the bill represents, which is a fair reading of a year in which the legal record grew faster than the legislative one.
What makes the timing worth noting is that the venue has changed. For two years the argument over machine access to the web was conducted in configuration files and vendor dashboards. That is where it stopped working. The clearest statement of why came from Cloudflare, and PPC Land reported it on September 27: "A robots.txt directive alone cannot solve this problem. Anyone can publish one, but it cannot identify who is crawling." The sentence is an admission from the largest operator of the relevant infrastructure that the honour system at the root of every domain is not a control at all. A robots.txt file states a preference to whoever chooses to read it.
Cloudflare's own settings changed on September 15, 2026, in a post written by Bryan Becker, its director of product security, under the title "Have it both ways: stay discoverable in search while disallowing AI training." A new Disallow AI Training option lets a site refuse model training while remaining indexed for search. At the same time the company extended its Block option to cover mixed-use crawlers operated by Google, Apple and Microsoft, the agents that fetch pages for search and for model work in the same pass. Legacy controls, including Block AI Bots and Managed Robots.txt, were deprecated, and new domains now receive presets based on whether they carry advertising. Customers were emailed on September 16 that automatic migration would run over the following week.
The practical consequence was stated plainly in the post: if mixed-use crawlers are to be gone entirely, a site now has to say so. That is a reversal of the guidance Cloudflare issued in July, when a Googlebot block was folded into the refusal of AI training by default. Two figures explain the retreat. About 17% of sites on Cloudflare enable some mechanism that blocks AI training. Fewer than 1% block search bots entirely. The second number is the constraint: almost nobody is willing to trade discoverability for refusal, so a setting that quietly imposed the trade was never going to hold. Cloudflare has moved through three positions in fifteen months, from pay per crawl to blocking opaque crawlers outright to paying per answer rather than per fetch.
The two tracks solve the same problem from opposite ends. A network setting asks the publisher to declare what it refuses, using directives such as Google-Extended, nosnippet and noarchive whose scope varies by vendor and by month. A statute would ask the crawler to declare what it is, which is the only half of the exchange that cannot be verified from the outside today. Publishers can see that GPTBot and ClaudeBot arrived because those agents announce themselves; the traffic that matters most is the traffic that does not. Trade bodies have begun writing their own taxonomies to fill the gap, with IAB Australia sorting every crawler into one of four verdicts. Meanwhile the underlying economics stay where they were: the New York Times spends about $2bn a year producing journalism that costs a model nothing to read.
Two dates now sit inside the same window. Microsoft has told publishers it aims to support robots.txt signals for its AI products in early 2027. Google's page-level generative search controls are scheduled for March 3, 2027, and the UK Competition and Markets Authority's publisher conduct requirement takes substantive effect on December 3, 2026. A bill that has been introduced but not passed is the slowest instrument in that list, and the only one that would apply to an operator who has never signed anything.
New Jersey decided publishers are data collectors
A second registry question surfaced the same week, and almost nobody affected by it appears to have noticed. AdExchanger's Allison Schiff reported on September 28 at 01:00 ET that New Jersey's data broker law reaches a category of company no other state statute targets: not the intermediary that buys and resells, but the collector that gathers consumer data directly and then sells or licenses it onward.
The law was introduced and passed within two days in June 2026 and took effect immediately on signing on June 30. Registration opens on April 1, 2027, and the fee schedule is steep at the top: $5,000 a year for a company selling or licensing the data of 100,000 or fewer New Jersey consumers, rising to $1.5m a year above 4.5 million. A separate prohibition has been live since the day of signature, barring the sale or licensing of health information, precise geolocation and biometric data outright.
The definitional point is the one that changes exposure. A data broker in most state regimes is a company with no direct relationship to the consumer. New Jersey's "data collector" is any company that collects consumer data directly and then sells or licenses it to even one data broker. Jodi Daniels, chief executive of Red Clover Advisors, put the consequence in six words: "There's no minimum thresholds." Many routine ad tech arrangements qualify, she noted, including arrangements where no money changes hands. Celine Guillou, special counsel at Kelley Drye, was blunter about the defence publishers have leaned on for three years: "Simply proclaiming 'We only use first-party data' is no longer an option." Daniel Rosenzweig, co-founder of DBR Tech Law, pointed to a cost that does not appear on the fee schedule, since qualifying publishers would appear on a public registry.
That last detail links this story to the one above it. A public list of companies that sell data is the mirror image of a bill requiring bots to name themselves. Both instruments work by compelled identification rather than by prohibition, and both put the administrative burden on whoever is easiest to find.
Enforcement is unsettled. Governor Mikie Sherrill's administration signalled in July 2026 that it planned to suspend enforcement while the legislature repaired the text, but the New Jersey Division of Consumer Affairs has confirmed that registration still opens on April 1, 2027, and the sensitive data prohibition remains enforceable whatever happens to the registry. A company that licenses health-adjacent segments through an identity graph or an enrichment partner is exposed now, not in 2027.
The pattern is familiar from other jurisdictions. California fined a data broker $45,000 for selling lists built on health conditions, and Healthline settled the largest CCPA case to date for $1.55m. What is different in New Jersey is that the statute does not ask whether a company thinks of itself as being in the data business. It asks what leaves the building. Congressional attempts to replace this patchwork with one federal standard, most recently the SECURE Data Act, have not advanced, so the count of separate registries keeps rising.
Six numbers where there used to be eight
If the first two stories concern who must declare themselves, the third concerns who decides what a number means. IAB Europe put version 2 of its In-Store Retail Media Definitions and Measurement Standards out for public comment on September 17, 2026, alongside a first draft Travel Media Measurement Addendum, and PPC Land examined the in-store document on September 27 at 18:23 UTC. The revision throws out the starting point the industry has used since the first version was finalised in December 2024. "Ad Play is not the right starting point for media measurement; footfall is," the draft states.
An impression becomes visual contact with an advertisement, and the draft builds a four-tier hierarchy beneath that definition: ad play, meaning the number of times an advertisement displays; gross impressions, meaning the individuals present while it displays; Opportunity To See, the best available proxy for a viewable impression; and Likelihood To See, which requires sensor technology and is flagged for its GDPR implications.
The calculation shrank from eight variables to six, a change agreed at an In-Store Measurement Workshop held on July 1, 2026. The worked example in the draft multiplies an average daily store footfall of 50,000 by 100 stores, 20 active days, three in-store placements, 80% aisle penetration and a 15% campaign share of voice, producing 36,000,000 Opportunity To See, or 0.36 per store visit. Three optional variables can be layered on for higher fidelity: average touchpoint compliance, OTS per visit and a location adjustment discount.
Aisle penetration is where the arithmetic meets the shop floor, and the draft ranks four ways of estimating it by precision: beacons in every store at aisle level; beacons in a representative sample extrapolated outward; inference from category transactions, treating a purchase as evidence of presence; and a third-party market study of category-level behaviour. Retailers must disclose which method produced their figure. That disclosure requirement extends further than in version 1, covering the footfall data source, the assumptions embedded in the calculation, whether a reported figure is OTS or LTS, aggregate or average dwell time, and a distance matrix of viewability by zone.
The draft is candid about what this means. "Without sensors, the assumptions behind the benchmark calculation (footfall, aisle penetration, touchpoint compliance) effectively are the measurement, which is why transparency on those assumptions is critical." Grocery illustrates the point: in grocery the starting point for footfall is often a transaction, and Tesco counts a purchase as an impression, de-duplicating where a shopper visits more than once inside the reporting window. Consumer electronics retailers cannot do that, because transaction data lags visits, so door scanners become essential. The five store zones from version 1 survive, running from external media in the car park through entrance areas, checkout, aisles and power displays to third-party services, click-and-collect and smart carts. Zone 1 borrows established out-of-home conventions. Zones 2 through 5 are described as assumption-driven without sensors.
Two changes cut the other way. The sales reporting window moved from 30-day pre and post periods to 28 days, four complete weeks, so that both windows carry an identical weekday distribution. And sales lift and brand sales lift were removed from the standard, a removal listed in the release announcement but absent from the draft's own summary table. Marie-Clare Puffett, IAB Europe's senior director of industry development and marketing, framed the exercise around confidence: "Measurement has always been fundamental to trust in digital advertising, and that's just as important as Commerce Media expands."
Other gaps are visible on a careful read. The OTS table still contemplates scenarios in which a displayed advertisement carries zero exposure, which sits awkwardly beside a definition built on visual contact. No independent audit mechanism is specified, despite IAB Europe running a certification programme. Programmatic integration is unaddressed, leaving a disconnect between footfall-based OTS and the per-play impression multipliers programmatic buyers already use. The contents page promises acknowledgments and an About section that the document does not contain.
The companion travel draft runs to nine pages, and PPC Land covered it on September 27 at 18:19 UTC. It was written with Booking.com, Expedia, Kayak, Skyscanner, Marriott Media, Koddi and Sojern. Sponsored product listings would carry a seven-day attribution window, display and offsite advertising 30 days, with last-touch attribution recommended and a three-year lookback for classifying a customer as new to an advertiser. The reporting metric proposed is Gross Bookings, counted at purchase and before cancellations, with base or total price reporting and mandatory fee disclosure. A travel network would therefore report a booking that a traveller cancelled the next morning as a sale, which is a defensible convention for comparing campaigns and a poor one for comparing revenue. Comment on both documents closes on October 23, 2026, at retailmediastandards@iabeurope.eu.
The context for the effort is a market measuring itself inconsistently while growing quickly. European retail media reached 13.7 billion euros in 2024, up 21.1%, with 20.8 billion euros forecast for 2026, and IAB Europe's own July 2026 AdEx Benchmark excludes in-store entirely. Earlier research found 70% of buyers naming an absence of standards as a barrier to investment. Version 1 was built with 14 retail media networks including Ahold Delhaize, Douglas Marketing Solutions, Kingfisher, MediaMarkt and Schwarz Media, and PPC Land covered that first release as well as the updated European definitions that followed. Consolidation has continued underneath the standards work: Perion agreed to acquire the in-store network PRN for up to $12m on August 25, 2026, covering 750 warehouse club, 4,500 big-box and 2,200 healthcare locations, and PRN helped write the eight-variable calculation that version 2 discards.
A merger ban that survived its own arithmetic
Booking.com helped draft the travel measurement standard. In the same week its parent lost a three-year fight over how much of the travel market it is allowed to hold. PPC Land reported the judgment on September 27 at 10:41 UTC: on September 9, 2026 the General Court of the European Union dismissed Booking Holdings' challenge to the European Commission's prohibition of its roughly 1.63 billion euro acquisition of Etraveli Group.
The case is T-1139/23, decided by the Tenth Chamber in extended formation, with President M. van der Woude sitting alongside M. Jaeger, L. Madise, P. Nihoul and S. Verschuur as rapporteur, after a hearing on July 8 and 9, 2025. The decision under challenge was C(2023) 6376 final in Case M.10615, adopted on September 25, 2023, at the end of an investigation that ran from a February 14, 2022 request by Booking for the Commission to take the case, through referral on March 9, notification on October 10, an in-depth investigation opened on November 16, a statement of objections on June 9, 2023, and two rounds of commitments in July and August of that year.
What makes the judgment interesting is that the court agreed with Booking about the numbers and upheld the ban anyway. Both of the Commission's quantitative methods for estimating the market share Booking would gain in flight distribution were found to contain errors. One applied no cannibalisation rate at all. There were attribution errors concerning what the Commission labelled Value Leadership and Halo Effects. Market size was held constant despite a projected 53% expansion. The baseline started from a zero-flights scenario the Commission itself had acknowledged was not the real one. Corrected, the share increase may amount to a few tenths of a per cent rather than the 0 to 5% band the decision relied on.
The court held that this did not matter, because of where Booking starts. In hotel online travel agency distribution across the EEA in 2022, Booking held 60 to 70%, with Expedia at 10 to 20% in business travel and 5 to 10% in consumer. In flights, eDreams Odigeo held 20 to 30%, Etraveli 10 to 20% and Booking 5 to 10%. Hotel dependence on the larger platform was the decisive fact: for most hotels, between 61% and 100% of bookings made through online agencies come via Booking, and 88% of independent hotels list on Booking against 61% on Expedia. In that structure, the judgment reasoned, "even a relatively small increase, in quantitative terms" can strengthen network effects, and the combined entity would create "a travel ecosystem which would be difficult for other OTAs to replicate." On cross-selling it found that flight data enables hotel targeting "not limited to the actual moment of a sale of the flight."
The Commission's own market investigation had produced an ambivalent picture that the court accepted as sufficient: 63% of hotels said they would not change the inventory they give Booking after the deal, 5% would reduce it, 27% would wait and see, and most expected Booking to take at least 5% from their direct channels. Booking's counter-arguments were rejected in turn, including a one-stop-shop efficiency claim ruled inadmissible because it post-dated the decision, an argument that channel management software makes multi-homing cheap, traffic data said to show minimal cross-selling potential, and the contention that rivals could equally use flights to build hotel customer bases.
Philipp Westerhoff of Geradin Partners in Berlin called the judgment particularly interesting for its treatment of digital market dynamics in ex ante merger control, and predicted it would become an important reference point beyond merger control. Booking has two months and ten days from notification to appeal to the Court of Justice, and stated that it was reviewing the judgment and a possible appeal. The record around it keeps thickening: the Court of Justice ruled on September 19, 2024 that Booking's parity clauses breached competition law, a Berlin court allowed 1,099 hotels to pursue damages on December 16, 2025, and the Commission fined Google 890 million euros on July 23, 2026 in a decision that named hotels among the affected verticals. Hotel groups and travel intermediaries have spent two years asking for stricter enforcement while documenting reduced visibility and greater reliance on intermediaries.
Whoever pressed the button owns the label
The last story returns to identification, this time of content rather than crawlers. IAB Austria's artificial intelligence working group, together with the law firm act legal Austria, published a five-page German-language memo titled "KI-Kennzeichnung im Digitalmarketing", and PPC Land analysed it on September 27 at 18:35 UTC. The document interprets the transparency obligations in Article 50 of EU Regulation 2024/1689, applicable since August 2, 2026, and carrying fines of up to 15 million euros or 3% of turnover.
Its central determination concerns the word deployer. Producing agencies are, as a rule, the responsible deployer, and must both assess whether a label is needed and implement it technically. Clients cross into deployer status when they specify the AI tools, issue prompting instructions, set parameters or operate their own systems. Publishers become deployers when they alter advertisements, personalise them, or embed them inside their own AI formats. The memo also marks a limit: "The mere knowledge that an advertising asset is being created with AI does not, taken on its own, establish deployer status." The working group was led by Iris Handlsberger of e-dialog, with Gerhard Gunther of DSR Agency, Maximilian Mondel of MOMENTUM, Stephanie Mauerer of e-dialog and Christoph Truppe, and the legal text was written by Mag. Philipp E. Stephan of act legal Austria.
Three triggers require a label: media that appears realistic, covering deepfakes, cloned voices, fabricated testimonials and manipulated places or events; AI-generated text on matters of public interest published without editorial review; and chatbots or interactive support tools. Several categories are exempt, including abstract AI graphics, cartoon characters and obviously fictional depictions, along with simple technical edits such as noise reduction, colour correction and retouching. Content created before August 2, 2026 is exempt unless it is republished or served again, which quietly converts a large archive into a compliance question the moment a campaign is reactivated.
On technique the guide is restrictive. Invisible watermarks, metadata and hidden imprint notes are insufficient. A label must be "clearly and understandably visible to users within the advertising asset itself or, for audio, audible." The EU icons released on June 10, 2026 are recommended, and the memo states that neither Google's SynthID watermark nor C2PA metadata satisfies the visibility test on its own. Two further deadlines sit behind that: machine-readable marking conformity on December 2, 2026, and watermark-detection interoperability on February 2, 2027.
The enforcement point is the part agencies will read twice. Austria's unfair competition law gives competitors and associations independent standing to seek injunctions and damages, entirely separately from any regulatory fine, and the memo states that "Both consequences can occur side by side." A missing label can therefore produce an administrative penalty and a civil claim from a rival for the same asset. Gaps remain: no enforcement authority is identified, obligations around emotion recognition and biometric categorisation are not covered, and no threshold is offered for the point at which a client's feedback rounds make it a deployer.
Austria is the second national trade body to allocate this duty in writing after a Dutch guide mapped four AI disclosure triggers for agencies, and the allocation matters because platforms have moved the other way. Google shifted AI ad labelling liability entirely onto advertisers in its own policies. The EU has published free labelling icons, publishers face the same 3% exposure, advertisers in New York face $5,000 penalties under a parallel state regime, and practitioners have argued that disclosure alone does not solve the underlying problem. There is also a measured cost to compliance: an NYU Stern study cited in the IAB's transparency framework found AI disclosure labels cutting click-through rates by 31.5%, a figure the Austrian memo repeats.
Four documents landed in three days, written by a Senate sponsor, a state legislature, a European trade body and an Austrian law firm, and every one of them is an attempt to make something declare what it is. A crawler, a data seller, an impression, a generated image. The work of the past two years went into building systems that move faster than anyone can inspect them. The work of this week went into writing down what they have to admit.
Also noted
- September 27: The Australian Competition and Consumer Commission collected $99,000 from the electronics retailer digiDirect over strikethrough discount claims, five infringement notices at $19,800 each, after finding that crossed-out prices such as $229 on a Logitech keyboard sold at $149 were rarely advertised and almost never charged.
- September 28: Delta Air Lines hired Keri Paison to build an advertising business, following United, whose inflight inventory reached programmatic buyers through Magnite.
- September 28: Google rewrote seven Merchant Center policy documents without changing the underlying rules, shifting the emphasis to account-level visibility restrictions and warning that a failed bulk appeal can limit later submissions.
- September 27: Amazon began letting US sellers connect eBay, Shopify, TikTok and Walmart accounts inside Seller Central, citing its own estimate that more than 95% of its independent sellers already sell across multiple channels.
- September 28: Adobe forecast $275bn in US holiday ecommerce spending despite consumer strain, a figure that will be tested from Prime Big Deal Days on October 6 and 7 onward.
By the numbers
- 17% Grocery order values climbed in the Americas while traffic fell 8%, on a Criteo survey of more than 6,300 consumers.
- 2.5x A Moulinex campaign run through Equativ and Valiuz beat its own first-quarter click-through rate, reaching 410,000 people between May 4 and June 7.
- $3.7bn The 650 films outside the top 20 took 43% of the 2025 US box office, a VAB argument against buying only tentpoles.
- 60 billion Product listings sit in Google's Shopping Graph, with 2 billion refreshed hourly, though AI Mode displays a fraction of them.
- October 23, 2026 Comment closes on the draft that redefines an in-store impression and on the travel addendum beside it.
Discussion