European websites that wrote a disallow line for AI fetching agents did not always get one. New measurement published on August 14, 2026 puts roughly 15% of identified page-fetching agents on the wrong side of that instruction, and the bot with the widest spread belongs to the most widely used chat assistant on the market.
The finding comes from the latest State of the Bots report by TollBit, the content licensing and bot monitoring company, covering the first half of 2026. It was reported on August 14, 2026 by Matt G. Southern at Search Engine Journal, whose account sets the report's numbers against the published crawler documentation of the companies operating those bots.
The measurement concerns a specific category of automated traffic. Not training crawlers, which sweep the web to build model corpora. Not search crawlers, which index pages so an answer engine can retrieve them later. The category at issue is the page fetcher: the agent that loads a single URL in real time because a person typed a question into a chat window and the assistant decided it needed that page to answer.
Where the disallow lines failed
Across the European sites discussed in the report, about 15% of identified AI page fetchers reached URLs that those sites had marked as disallowed, according to TollBit's data as described by Search Engine Journal.
That aggregate figure hides an uneven distribution. The bypasses concentrate in a handful of named agents. ChatGPT-User, Bytespider and Youbot each accessed disallowed pages on nearly half of the European sites that had explicitly listed them in robots.txt. Among those three, ChatGPT-User reached the largest number of sites.
The same agent holds a second distinction. ChatGPT's page-fetching bot is disallowed by more sites than any other bot of its type, and it also reached disallowed pages on more sites than any other bot. Being the most frequently named agent in robots.txt files and the agent most often observed on restricted URLs is not a contradiction; it is the arithmetic of an instruction that carries no enforcement.
TollBit's methodology is deliberately blunt on this point. According to Search Engine Journal's report, the company counts any request to a disallowed URL as a bypass, without regard to what the operator's documentation claims about its own compliance.
Europe blocks the newer agents far less than North America
A second set of figures in the report describes what site owners are actually writing into their files, and here the regional split is wide.
Only 9% of European websites disallow Claude-User, the fetching agent operated by Anthropic, compared with 26% in North America. Perplexity-User shows a similar pattern at 13% in Europe against 26% in North America. Most of the newest page-fetching agents sit in single-digit disallow rates across Europe.
ChatGPT-User is the exception. It is blocked at rates that put it well clear of its peers, which is why it also produces the largest absolute count of observed bypasses: an agent nobody has named in robots.txt cannot register a violation of a rule that was never written.
That statistical artefact matters for interpretation. A low bypass count for a newer agent does not establish good behaviour. It may only establish obscurity.
What the operators say in their own documentation
The reason a disallow line for ChatGPT-User does not carry the weight a publisher might assume is written into OpenAI's own developer documentation. The company states that ChatGPT-User visits a page when a ChatGPT user asks a question, and that because such actions are initiated by a user, robots.txt rules may not apply.
Perplexity takes the same position for Perplexity-User. Anthropic does not, and states that all three of its bots respect the file, a position PPC Land documented when the company clarified what ClaudeBot, Claude-User and Claude-SearchBot each collect in February 2026.
The divergence is not new. OpenAI removed robots.txt compliance language for ChatGPT-User from its crawler documentation on December 9, 2025, a change spotted by consultant Pieter Serraris. Google formalised the same category of traffic on March 20, 2026, when it added Google-Agent to its official list of user-triggered fetchers, documenting that such fetchers generally ignore robots.txt because a person requested the page. IAB Australia later built the disagreement into formal guidance, sorting every crawler into one of four verdicts and concluding that robots.txt is not a reliable control for live agents.
Three of the largest operators of assistant traffic now hold two incompatible positions between them on whether the same file governs the same behaviour.
The visibility half of the trade
OpenAI's crawler overview separates its user agents by function, and the separation carries consequences that are easy to miss when a site owner blocks everything with the letters AI attached.
According to the documentation, OAI-SearchBot is the agent used to surface websites in search results within ChatGPT's search features. Sites opted out of OAI-SearchBot will not appear in ChatGPT search answers, though they can still show up as navigational links. OpenAI recommends allowing that agent in robots.txt, along with requests from its published IP ranges, for sites that want to appear in those results.
Each setting operates independently. A site can allow OAI-SearchBot to appear in search results while disallowing GPTBot to signal that crawled content should not be used for training OpenAI's foundation models. Where both are allowed, OpenAI states it may use the results of a single crawl for both purposes to avoid duplicate crawling. Changes take time to register: the documentation notes that search systems can require roughly 24 hours from a robots.txt update to adjust.
The asymmetry follows directly. A site that disallows both OAI-SearchBot and ChatGPT-User has given up the search visibility, which OpenAI documents as an honoured instruction, and retained a fetching control that OpenAI documents as carrying a carve-out. As Search Engine Journal framed it, sites blocking both agents "have traded away the visibility half of that deal."
Existing research complicates any assumption that blocking is cost-free. Rutgers Business School and The Wharton School found that news publishers who blocked AI crawlers through robots.txt lost roughly 7% of weekly website traffic within six weeks, a decline visible in human browsing panel data rather than bot metrics alone. Separate analysis found that blocking does not reliably remove a site from AI citation datasets either.
Enforcement moves to the network layer
The response taking shape sits below robots.txt entirely.
Cloudflare published a policy update on July 1, 2026 titled "Your site, your rules: new AI traffic options for all customers," written by Jin-Hee Lee and Bryan Becker. The post replaces the company's earlier binary framing of AI bots with a taxonomy built on behaviour rather than label. Three categories anchor it: Search, covering behaviour that collects or indexes content to answer questions about it later; Agent, covering automated behaviour acting in real time on a person's behalf; and Training, covering crawlers taking content to train or fine-tune a model.
The Agent category names its examples directly. According to Cloudflare, it includes chat fetch bots such as ChatGPT-User and browser-use agents such as Gemini or Claude driving Chrome. The bots at the centre of TollBit's bypass measurement are, under this scheme, a distinct class with their own controls.
Cloudflare's stated rationale for splitting the categories is that locking content down is not, in the company's phrasing, a matter of website owners resorting to "block all automation, every time." The company also argues that operators running search indexing, agent activity and training collection should separate that automation into three crawlers rather than one.
The September 15 defaults
The dated element of the policy is a change to defaults. From September 15, 2026, for all new domains onboarding to Cloudflare, crawlers classified as Training and Agent will be blocked by default on pages that display ads. Search will remain allowed by default.
The reasoning Cloudflare gives ties the rule to monetisation. An advertisement, the company argues, signals that a website owner intended a person to land on the page and see it, so on those pages human attention is treated as the objective and bots that may displace it are kept away.
A second change lands the same day. Multi-purpose crawlers that combine Search with Training will be allowed or blocked according to all of their behaviours, with defaults enforced by the most restrictive applicable rule. Cloudflare names Googlebot, Applebot and BingBot as examples that will be blocked for customers who have selected to block Training, whether through the new controls or the legacy Block AI bots service. Existing customers can mark an opt-out in Security settings at any point before that date.
PPC Land covered the companion announcements from the same day, including Cloudflare's shift from charging AI companies per crawl toward paying publishers based on whether content was used to generate an answer and the Attribution Business Insights dashboard that put crawl-to-referral ratios in front of publishers for the first time.
BotBase, content use and transitive trust
Three further mechanisms in the July post extend the same logic.
BotBase is a new searchable database of known bots and agents available to Enterprise Bot Management customers on the Cloudflare dashboard, showing where each tracked bot sits in the updated taxonomy and exposing a detection ID that can be copied into security rules. The classification list runs well beyond the three headline categories, covering Transact, Data Collection, Security Testing, SEO, Ads Verification, Social and Link Preview, Feed Fetching, and Monitoring and Operations.
Alongside classification, Cloudflare is building controls based on content use, meaning what a bot keeps and reshares after access. Three levels apply, from least to most permissive: immediate, meaning interact but store and reuse nothing; reference, the default, meaning index, excerpt and link back; and full, meaning summarise and reproduce. A matching robots.txt signal called use extends the Content Signals specification with a fourth optional field, and customers already running Cloudflare's managed robots.txt have use=reference added to theirs. Bots that reproduce content in full cannot hold Verified status, and bots found abusing the signals lose it.
The definition of Verified itself narrows. Previously all Verified bots were allowed by default. Under the new arrangement, non-verified bots remain blocked by default while the Verified label only makes a bot allowable within its relevant category, so the categories a site owner permits determine access.
The final proposal addresses a structural problem in agent traffic: the bot at the door is often not run by the company that built the tooling. Cloudflare describes this as transitive trust and proposes using the existing Forwarded header defined in RFC 7239 to carry operator identity and declared content use through intermediaries, in a format such as Forwarded: for="openai";use="reference". The enforcement argument rests on scale. Losing trusted status across the more than 20% of web domains behind Cloudflare, the company writes, is "a deterrent with teeth."
Why this matters for marketers and publishers
The gap between what a robots.txt file requests and what a server log records is now a measurable quantity rather than a suspicion, and the measurement has commercial consequences on both sides of the media transaction.
For publishers running advertising, the September 15 default sits at the intersection of the two stories. Agent traffic is precisely the class TollBit measured bypassing disallow lines, and it is precisely the class Cloudflare will block by default on ad-carrying pages for new domains. A control that operators can decline to honour is being replaced, for a growing share of the web, by one they cannot.
For media buyers and measurement teams, the composition of traffic is the practical issue. Automated requests already exceed human ones on Cloudflare's network, and PPC Land has documented the consequences across the stack, from AI bots hitting WordPress cart pages 3.75 million times in a day to Microsoft Clarity surfacing robots.txt violations inside Bot Analytics in June 2026. Traffic that never had a person attached distorts inventory counts, session metrics and the denominators in every conversion rate a campaign reports.
For European operations specifically, the disallow rates in the TollBit data describe a market that has written fewer rules for the newest agents than North America has. Whether that reflects deliberate openness, slower documentation cycles or simple unawareness is not something the bypass numbers can settle.
The unresolved question is the one Search Engine Journal ends on: whether the user-initiated carve-out survives at all. It rests on an argument that a request made because a person asked differs in kind from a crawler taking content on its own schedule. Every major assistant now fetches pages this way, and the volume behind that distinction keeps rising.
Timeline
- December 9, 2025 - OpenAI revises its ChatGPT crawler documentation, removing robots.txt compliance language for ChatGPT-User in user-initiated browsing
- February 25, 2026 - Anthropic clarifies that ClaudeBot, Claude-User and Claude-SearchBot all respect robots.txt
- March 20, 2026 - Google adds Google-Agent to its official list of user-triggered fetchers, documenting that such fetchers generally ignore robots.txt
- April 21, 2026 - Rutgers and Wharton researchers post the revised study finding a 7% weekly traffic decline for publishers that blocked AI crawlers
- June 23, 2026 - Microsoft Clarity adds robots.txt violation detection to Bot Analytics
- July 1, 2026 - Cloudflare publishes "Your site, your rules," introducing Search, Agent and Training controls for all customers, and separately shifts from per-crawl charging to answer-based compensation
- July 31, 2026 - IAB Australia issues crawler guidance concluding robots.txt is not a reliable control for live agents
- August 14, 2026 - Search Engine Journal reports TollBit's State of the Bots findings for the first half of 2026, including the 15% European bypass rate
- September 15, 2026 - Cloudflare's new defaults take effect, blocking Training and Agent crawlers on ad-carrying pages for newly onboarded domains
Related PPC Land coverage
- OpenAI revises ChatGPT crawler documentation with significant policy changes - The December 2025 documentation change that removed robots.txt compliance language for ChatGPT-User.
- Anthropic clarifies what its three web crawlers do - and how to block them - Anthropic's position that all three of its bots honour robots.txt, in contrast with OpenAI and Perplexity.
- Google-Agent joins the crawler list as AI browsing gets an official identity - Google's March 2026 formalisation of user-triggered fetchers that bypass robots.txt by design.
- Cloudflare stops charging AI per crawl and starts paying per answer - The companion July 1 announcement on compensation, with the September 15 default-blocking date.
- Cloudflare exposes AI crawlers hitting sites 50000 times per visitor - The Attribution Business Insights dashboard and its Training, Search and Agent classification.
- Microsoft Clarity now flags robots.txt violations inside Bot Analytics - Measurement tooling that shows which crawlers ignore access rules and what they target.
- IAB Australia forces every crawler into one of four verdicts - Industry guidance classifying live agents separately and recommending edge enforcement.
- Blocking AI crawlers cost news publishers 7% of traffic, study finds - Academic evidence on the traffic cost of robots.txt blocking.
- Blocking AI crawlers doesn't stop citations - new data shows why - Evidence that blocked sites still appear in AI citation datasets.
- The user agent strings every SEO and site owner needs right now - Reference on OpenAI, Anthropic and Google identifiers and their share of observed AI traffic.
- llms.txt adoption rises 8.8x but 97% of files get zero AI requests - Data on the gap between publishing machine-readable preferences and having them read.
- AI bots hammered WordPress cart pages 3.75M times in a day, Kinsta data shows - Infrastructure and commerce effects of automated traffic at scale.
Summary
Who: TollBit, which produces the State of the Bots report; OpenAI, Anthropic, Perplexity and ByteDance, whose page-fetching agents appear in the measurement; Cloudflare, whose July 1 policy moves crawler decisions to the network layer; and Matt G. Southern of Search Engine Journal, who reported the findings.
What: Roughly 15% of identified AI page fetchers reached URLs that European sites had marked as disallowed. ChatGPT-User, Bytespider and Youbot each reached disallowed pages on nearly half of the European sites naming them, with ChatGPT-User covering the most sites. Only 9% of European sites disallow Claude-User against 26% in North America, and 13% disallow Perplexity-User against 26%. OpenAI's documentation states robots.txt may not apply to ChatGPT-User because a person initiated the request.
When: The report covers the first half of 2026 and was reported on August 14, 2026. Cloudflare's policy update was published on July 1, 2026, with new defaults taking effect on September 15, 2026.
Where: The bypass and disallow figures describe European and North American websites. Cloudflare's default change applies to new domains onboarding to its network, which covers more than 20% of web domains.
Why: Robots.txt is a request rather than an enforcement mechanism, and several operators of user-triggered fetching agents state in their own documentation that it may not govern their behaviour. That gap has commercial weight for publishers whose ad-supported pages absorb the traffic, for advertisers whose metrics include it, and for anyone whose search visibility depends on which agents a site chooses to admit.
Discussion