Cloudflare on August 21, 2026 published Bot Preference Sync, a feature that rewrites a website's robots.txt file to match the AI crawler policy already configured in its dashboard, and attached four disclosure conditions that determine whether a crawler doing both search and training keeps access to sites refusing training.
Two files have governed how a website states its position on automated traffic, and for years they have been permitted to contradict each other. One is robots.txt, the plain-text declaration a crawler reads before fetching a page. The other is the enforcement layer: firewall rules, bot management settings and edge blocks that act regardless of what any crawler claims to honour. Cloudflare argues the gap between the two has become a liability for the site owner rather than for the crawler.
The company's post, written by Jin-Hee Lee and tagged under AI Bots, Bot Management and Product News, sets out the failure mode without ornament. A robots.txt file can state that a crawler is disallowed while the enforcement rules quietly let that same crawler through. When the stated preference and the enforced rule disagree, according to Cloudflare, some crawlers treat the inconsistency as a basis to disregard the preference or to attempt a bypass of the enforced rules.
Bot Preference Sync is the response. Instead of a site owner maintaining a separate static file by hand, policy is configured once at the zone level and Cloudflare generates or updates robots.txt from that configuration. The feature is stated as available to all customers, from the free tier to Enterprise, and can be turned on or off at any time. Availability was given as the coming week, with the company's changelog named as the place where the rollout will be confirmed.
How the file gets written
The mechanism is additive rather than destructive. Where a site already has a robots.txt file, the contents generated by Bot Preference Sync are prepended to the existing material, so any Disallow directives already present survive the operation. The generated section is fenced by comment markers. The example published in the post, shortened and anonymised by Cloudflare, opens with a line reading BEGIN Cloudflare Bot Preference Sync, lists four user agents named TrainingBot1, TrainingBot2, TrainingBot3 and MixedUseBot-Extended, applies a site-wide Disallow, and closes with a matching END marker.
The membership of that list is not fixed at the moment of configuration. Cloudflare states that it will draw on the bots it tracks in BotBase to periodically refresh the set of user agents written into robots.txt whenever a customer chooses to block or disallow a given category. Verified bots classified as Search, Agent and Training can be inspected at any point through the company's public bots directory. The practical effect is that a preference expressed once against a category continues to cover new crawlers entering that category, without the site owner editing anything.
That design carries a stated limit. Bot Preference Sync addresses policy decisions taken across a whole category rather than negotiated crawler by crawler, and it will not read from individual custom rules containing more complex logic. A customer that has struck a particular arrangement with a named company, and wants an exception carved for that company alone, retains the option of switching the sync off and hand-tailoring the file to match the custom policy.
Three categories, and a narrower meaning for Disallow
The taxonomy underneath the feature is the one Cloudflare introduced on July 1, 2026, when it replaced its earlier binary framing of AI bots with three behavioural categories: Search, covering collection and indexing of content to answer questions about it later; Agent, covering automation acting in real time on a person's behalf; and Training, covering crawlers gathering material to train or fine-tune a model.
For Search and Agent, the three settings announced in July remain unchanged: allow, block on pages that serve ads, or block everywhere. Training is where the August 21 post makes a substantive change. The Disallow option now writes a no-training preference into robots.txt in a form that lets cooperating mixed-purpose crawlers continue to reach the content for search indexing. The reasoning given is conditional rather than charitable: those crawlers keep access because they permit site owners to verify directly how the data is used. Search visibility for cooperating crawlers, according to Cloudflare, is unaffected.
The company's position on the underlying problem has not softened. Its July post argued that mixed-use crawlers, described as "bots that blend search, agent use, and training behind a single user agent," place site owners at a disadvantage precisely because they make it difficult to separate wanted behaviour from unwanted behaviour. That argument is restated as still standing.
Transparency becomes a condition of verification
What is new is the price attached. For the purposes of bot Verification, operators whose bots perform both Search and Training must supply additional information in order to avoid being blocked on sites where Disallow Training has been set. Four requirements are listed.
First, the bot must respect a no-training preference in robots.txt, through any mechanism. Second, the operator must give site owners a route to opt out of AI summaries. Third, the operator must provide URL-level visibility into which pages were made available for training, together with metrics on search results, so that a site owner can see how content was used on each side. Fourth, the operator must be able to demonstrate publicly that disallowing training does not damage traditional search results.
Bots from leading AI model developers and service providers that satisfy those criteria are to be tracked publicly in the AI bot transparency section of Cloudflare Radar, which the company says will include cases where best practices are honoured alongside cases where they are not. Crawlers that decline to supply the disclosure receive no benefit of the doubt and remain blocked wherever training is disallowed. Cloudflare frames the arrangement as "making Transparency the price of admission."
The four conditions are not abstract. Each maps onto a dispute already on the record. The second and fourth in particular describe the exact controls publishers have spent two years demanding from the largest search operator. Google described building a separate opt-out for AI Overviews without loss of search visibility as a major engineering project in February 2026. A control appeared on June 3, 2026, the same day the United Kingdom's Competition and Markets Authority imposed its first binding Publisher Conduct Requirement under the Digital Markets, Competition and Consumers Act 2024. That control operates at domain level rather than page level, with page-level grounding scheduled for March 3, 2027 and the substantive obligations of the requirement taking legal force on December 3, 2026.
The third requirement, URL-level reporting on what was taken for training, has no equivalent at any major operator. Anthropic separated the functions of ClaudeBot, Claude-User and Claude-SearchBot in documentation published on February 25, 2026 and committed all three to honouring robots.txt, which addresses the first condition but not the reporting ones. OpenAI moved in the opposite direction in December 2025, when it removed robots.txt compliance language from the ChatGPT-User description on the reasoning that user-initiated fetches are not automated crawling.
A separate starting position for ad-supported sites
The post also changes what a new domain inherits on day one, and it splits the default in two.
At onboarding, a customer can select an option stating that the domain monetises from pages carrying advertising. Selecting it sets Training to Disallow as the default, on the stated expectation that sites depending on advertising revenue reserve those pages for human visitors. The setting is changeable at any later point. Cloudflare summarises the intended outcome as remaining in search while keeping content out of model training.
For every other new customer, nothing is applied. No blocks and no disallows are added at onboarding, and the starting point adds no restrictions on the customer's behalf. Search, Agent and Training can each be blocked afterwards at the customer's choosing.
For all new customers, Bot Preference Sync itself is switched on by default. Existing customers running the legacy managed robots.txt feature will be prompted to review and confirm their preferences in order to transition to the new system when it launches. That legacy feature dates to the earlier managed robots.txt work, which paired a managed file telling a fixed list of major training crawlers not to train on a site's content with edge-enforced blocks against those same crawlers.
What sits underneath the announcement
The numbers driving the category are not in this post, but they are on the record. Cloudflare's Attribution Business Insights dashboard, opened on July 1, 2026, documented crawl-to-referral ratios running from 118 crawls per referral at the low end to nearly 50,000 at the high end. By early June 2026, bots accounted for 57.4% of all traffic to HTML content across the network, with training-related crawlers alone at 50.6% and search crawlers at 10.7%. On the same July date, the company abandoned per-crawl charging in favour of paying publishers when content contributes to a generated answer, and set September 15, 2026 as the date from which Training and Agent crawlers are blocked by default on advertising-carrying pages for domains newly joining the network.
Compliance with stated preferences remains partial. TollBit measurement reported in August 2026 found 15% of AI page fetchers in Europe reaching URLs that had been disallowed. Microsoft Clarity added robots.txt violation flags inside its Bot Analytics view in June 2026, supplying evidence without enforcement. Cloudflare itself de-listed Perplexity as a verified bot in 2025 after documenting undeclared crawlers using a generic browser user agent, and the enforcement route it built in December 2024, Robotcop, translated robots.txt rules into firewall rules applied at the network edge.
Refusal has a measured cost on the other side of the ledger. Research from Rutgers Business School and The Wharton School found news publishers that blocked AI crawlers through robots.txt lost roughly 7% of weekly website traffic within six weeks, a figure reported elsewhere in the same body of work as 23.1% of monthly visits, without proportional protection, because the protocol remains voluntary. That asymmetry is the reason the Disallow refinement matters: it is an attempt to make the training refusal separable from the search penalty.
Why this matters for advertising and publishing
The advertising layer is now explicitly written into infrastructure policy. Cloudflare has drawn its default line around pages that carry advertising twice in eight weeks, first with the September 15 blocking rule and now with an onboarding question that assigns ad-funded domains a different training default from everyone else. An advertisement, in the reasoning Cloudflare has offered, functions as a declaration that the page was built for a person to see.
For publishers running programmatic inventory, the operational consequence is that a robots.txt file stops being a document maintained by a technical team and becomes an output of a commercial policy setting. The two artefacts that auditors, licensing counterparties and AI operators inspect will, by construction, say the same thing. Sites that have quietly relied on the discrepancy between a permissive robots.txt and restrictive edge rules, or the reverse, lose that ambiguity unless they switch the sync off deliberately.
For measurement teams, the third disclosure requirement is the one with the longest reach. URL-level visibility into which pages were made available for training, paired with metrics on search results, would give a publisher something no current dashboard provides: a per-page account of what was taken and what came back. Whether any large operator supplies it is the open question, and Cloudflare Radar's transparency section is where the answer becomes public.
For the AI operators, the calculation is now positional. Cloudflare sits in front of more than 20% of the web, a share that turns a verification policy into a distribution question. A mixed-purpose crawler that meets the four conditions retains search access across sites that have disallowed training. One that does not is blocked on those same sites, and the blocking is applied at the edge rather than requested in a text file. Whether that is enough pressure to change operator behaviour, on a network where roughly 15% of European fetchers already ignore what robots.txt says, is the test the coming months will run.
Timeline
- December 10, 2024 - Cloudflare launches Robotcop, converting robots.txt rules into network-level firewall enforcement
- July 1, 2025 - Cloudflare opens pay per crawl in private beta using HTTP 402 responses, alongside a managed robots.txt option covering major training crawlers
- August 28, 2025 - AI Crawl Control expands with customisable 402 responses
- December 9, 2025 - OpenAI removes robots.txt compliance language from the ChatGPT-User description
- February 11, 2026 - A Google executive describes a separate AI Overviews opt-out as a major engineering project
- February 25, 2026 - Anthropic documents the separate roles of its three crawlers and commits them to robots.txt
- June 3, 2026 - The CMA imposes its first binding Publisher Conduct Requirement and Google begins testing a Search Console opt-out toggle
- June 23, 2026 - Microsoft Clarity begins flagging robots.txt violations inside Bot Analytics
- July 1, 2026 - Cloudflare introduces the Search, Agent and Training categories, opens the Attribution Business Insights dashboard and shifts from per-crawl charging to answer-based compensation
- August 2026 - TollBit measurement finds 15% of AI page fetchers in Europe reaching disallowed URLs
- August 21, 2026 - Cloudflare publishes Bot Preference Sync, the four transparency conditions for mixed-purpose bot verification, and the separate training default for ad-monetised domains
- Week of August 24, 2026 - Stated availability window for Bot Preference Sync across all plans
- September 15, 2026 - Default blocking of Training and Agent crawlers on advertising-carrying pages begins for domains newly joining the network
- December 3, 2026 - Substantive obligations of the CMA Publisher Conduct Requirement take legal force
- March 3, 2027 - Page-level grounding controls scheduled for publishers
Related PPC Land coverage
- Cloudflare exposes AI crawlers hitting sites 50000 times per visitor - Details the Attribution Business Insights dashboard and the Training, Search and Agent classification it applies per operator.
- Cloudflare stops charging AI per crawl and starts paying per answer - Covers the July 1, 2026 pricing reversal and the September 15 default-blocking date for ad-carrying pages.
- Cloudflare ties AI payouts to citations as 50% of crawls waste - Sets out the traffic composition figures, including bots at 57.4% of HTML traffic in early June 2026.
- 15% of AI page fetchers in Europe reached disallowed URLs, TollBit finds - Measures the gap between stated robots.txt preferences and actual crawler behaviour in Europe.
- Cloudflare launches Robotcop to enforce robots.txt policies against AI crawlers - The December 2024 origin of edge enforcement built directly on robots.txt directives.
- Anthropic clarifies what its three web crawlers do - and how to block them - Documentation separating ClaudeBot, Claude-User and Claude-SearchBot, with all three committed to robots.txt.
- OpenAI revises ChatGPT crawler documentation with significant policy changes - The December 2025 change removing robots.txt compliance language for user-initiated fetches.
- Perplexity denies training AI models as Cloudflare documents stealth crawlers - Cloudflare's de-listing of a verified bot after documenting undeclared crawling activity.
- UK regulator forces Google to give publishers AI opt-out rights today - The CMA conduct requirement behind the AI summary opt-out that Cloudflare now lists as a verification condition.
- Google gives site owners a toggle to exit AI Overviews and AI Mode - The domain-level scope of that control and its March 2027 page-level deadline.
- Microsoft Clarity now flags robots.txt violations inside Bot Analytics - Measurement tooling that surfaces which crawlers ignore stated access rules.
- US sends 53.5% of global bot traffic, Decodo analysis finds - Wider bot traffic composition and the documented cost of blocking through robots.txt.
- IAB Australia forces every crawler into one of four verdicts - Industry guidance concluding that robots.txt is not a reliable control for live agents.
- Cloudflare and ETH Zurich say AI bots are breaking the web's cache layer - Joint research on how automated traffic defeats caching assumptions built for human browsing.
Summary
Who: Cloudflare, whose network sits in front of more than 20% of the web, through a blog post written by Jin-Hee Lee. The affected parties are site owners on every Cloudflare plan, publishers monetising through advertising, and the operators of AI crawlers that combine search and training behaviour under a single identity.
What: Bot Preference Sync, a feature that generates or updates a site's robots.txt file from the AI bot policy already configured in the Cloudflare dashboard, prepending its directives to any existing file. Alongside it, four disclosure requirements now determine whether a mixed-purpose crawler avoids being blocked on sites that disallow training, and ad-monetised domains receive Disallow as their training default at onboarding.
When: Published August 21, 2026, with availability stated for the following week across all plans. The taxonomy it builds on dates to July 1, 2026, and the related default blocking of Training and Agent crawlers on advertising pages begins September 15, 2026.
Where: Across Cloudflare's global network, configured at the zone level in the dashboard, with verified bot classifications published in the company's public bots directory and compliance behaviour tracked in the AI bot transparency section of Cloudflare Radar.
Why: Divergence between what a site declares in robots.txt and what it enforces at the edge has been treated by some crawlers as grounds to ignore the declaration. Synchronising the two removes that argument, while the verification conditions attempt to make disclosure a precondition of continued search access rather than a voluntary courtesy.
Discussion