HasData, a web scraping API provider, updated its AI Crawler Block Index on September 29, 2026, adding a September 16 re-run of 10,894 websites that shows Cloudflare's September 15 switch stripping machine-readable no-training lines from the files of sites that carried them. The July baseline had already found that two in five sites banning OpenAI's crawler in their robots.txt files still served it pages.
In Short
A web scraping company tested about 10,900 websites in July 2026 and again in September to see which ones really keep AI robots out, and found that many sites say one thing in their public rules file and do another at the door. That matters to publishers and advertisers because a Cloudflare change on September 15 wiped the AI rules from the files of 118 sites that carried them, which makes the public file a weaker guide to what a site actually does. In practice, what a site announces and what its server does can differ, so any count of who blocks AI depends on whether the file or the server was measured.
What HasData measured
According to HasData, the study covers 10,894 registrable domains: 9,746 drawn from the Tranco top-10,000 list (list 46XLX, dated July 9, 2026) and 1,148 news homepages from a set compiled by Ben Welsh, after removing duplicates. Scope and declared-policy questions ran against all of them; enforcement questions ran against a 2,096-domain subset of the top 1,000 open-web sites plus every publisher.
Each domain in the subset received two requests from a US datacenter IP address pinned to the same exit, one announcing itself as OpenAI's GPTBot and one as a standard Chrome browser. The browser acts as a control: if it gets a page and the crawler does not, identity is the variable. A 600-domain sample also received a residential-IP pass. About 15% of enforcement requests hit proxy errors and were excluded, while another 8.2% met a challenge or CAPTCHA and were reported separately.
Several definitions carry the counts. A paper-only ban is a site whose robots.txt disallows GPTBot but whose server returned a live 200 to a GPTBot request. A silent enforcer is the reverse: robots.txt says nothing, yet GPTBot is blocked while the browser control receives a 200. The IP gap counts domains that block a datacenter address but serve a residential one. A bot counts as blocked in robots.txt only when its own user-agent group contains "Disallow: /" with no overriding Allow, and a site counts as behind Cloudflare when its headers include cf-ray or server: cloudflare.
HasData states three limits. The September comparison is change over time rather than a controlled experiment, and the one cause it names comes from Cloudflare's own announcement. Its GPTBot requests carry the official user-agent string but are not OpenAI's verified crawler, so a site matching on verified identity would treat them differently. The AI Mode test and the five-crawler residential test were not repeated.
HasData sells a web scraping API, and the author, co-founder and chief technology officer Roman Milyushkevich, describes his work as proxy infrastructure and data extraction.
Publishers close, the open web mostly does not
The July baseline split the sample into two very different populations. According to HasData, 10.3% of the top web blocks at least one AI crawler in robots.txt, against 56.4% of news publishers - a ratio of 5.5. For GPTBot alone the figures are 7.9% and 50.5%. ClaudeBot, CCBot and Google-Extended show the same pattern at similar magnitudes, the document says.
| Metric | Top web | News publishers |
|---|---|---|
| Behind Cloudflare's proxy | 26.0% | 22.4% |
| In scope (Cloudflare plus ad-tech signals) | 8.5% | 13.6% |
| Blocks any AI crawler in robots.txt | 10.3% | 56.4% |
| Blocks GPTBot in robots.txt | 7.9% | 50.5% |
| Paper-only bans | 46.7% | 38.9% |
| GPTBot served / browser served (same datacenter IP) | 67.8% / 72.3% | 54.2% / 83.8% |
| Publishes llms.txt | 7.9% | 3.2% |
HasData attributes the gap to business model. Publishers hold text that models want to train on and cite, and carry subscription revenue that competes with AI summaries, while most e-commerce, blog and documentation sites do not see AI training as a direct risk. In the page's formulation, "news publishers are closing to AI" is the accurate version of any headline claiming the whole web is doing so.
Bans on paper, enforcement elsewhere
Of 592 sites in the enforcement subset that disallow GPTBot, 234 (39.5%) still served the crawler a live page, according to HasData. In the author's words, "robots.txt is a suggestion". Open-web sites ignored their own file more often (46.7%) than publishers (38.9%), though the open-web figure rests on only 44 to 45 sites, which HasData calls too small a base to carry a headline.
Because the requests were not OpenAI's verified crawler, a 200 may reflect a server that checks identity more closely than a user-agent string. Independent evidence on compliance points the same direction: a TollBit report found that about 15% of identified AI page fetchers in Europe reached disallowed URLs, and the same coverage notes OpenAI's position that robots.txt may not govern ChatGPT-User. OpenAI's documentation keeps GPTBot and OAI-SearchBot under robots.txt, as a June 28 guide to crawler user-agent strings lays out. Whether a bot honours the file is a different question from whether a server stops it, and the index measures only the second.
When a site did enforce something against GPTBot, HasData lists 22.2% hard blocks (HTTP 403, 401 or 451), 10.9% HTTP 402 Payment Required, 5.0% Cloudflare-style challenges or JavaScript interstitials, 2.4% rate limits (429) and 2.3% CAPTCHAs. The page does not state the denominator for these shares. The 402 line ties to Cloudflare's pay per crawl programme, which entered private beta on July 1, 2025 and uses that code to signal that a price applies; for a bot without a paid arrangement, HasData notes, the outcome matches a hard block.
The reverse case is the silent enforcer. In July, 115 sites (5.5% of the subset) blocked GPTBot without disallowing it in robots.txt, filtering at the firewall instead; by September the count was 178, made up of 45 open-web sites and 133 publishers. HasData's phrase for the pattern is that declared policy and behavior "diverge in both directions." For a publisher, that means the public file can understate the protection in place; for a crawler operator, it means the file gives no warning of the block ahead.
An identity penalty
From one datacenter IP, publisher sites served the Chrome request 83.8% of the time and the GPTBot request 54.2% - a gap of roughly 30 points. Whole-web sites showed the same direction with less spread: 72.3% against 67.8%. HasData concludes that enforcement is identity-based, since sites match on the User-Agent header rather than on IP ranges belonging to OpenAI.
A five-crawler test in July, run from a residential address against 250 publisher domains, extended the result beyond one bot.
| User agent | Pass rate | Drop from browser |
|---|---|---|
| Browser (Chrome 150) | 77.2% | baseline |
| GPTBot | 44.4% | 33 points |
| ClaudeBot | 36.8% | 40 points |
| PerplexityBot | 45.2% | 32 points |
| CCBot | 37.6% | 40 points |
| Meta-ExternalAgent | 38.0% | 39 points |
HasData's explanations for the ordering, such as publisher licensing deals with OpenAI, are interpretation; the page presents no test of them. The document also contains a discrepancy: the body describes a gap of about 30 points on a datacenter IP, while the FAQ says a browser user-agent was served "20 to 30 points higher". The tables support the larger figure.
Rotating IP addresses did little. Of 450 domains that blocked GPTBot from the datacenter address, 48 (10.7%) served it from the residential one; the September re-run puts that share at 6.8%. If identity carries the load, a crawler that omits or misstates its user agent meets a lighter gate - a pattern PPC Land documented in January 2026, when one Grok query triggered 16 requests from 12 IP addresses, none identifying the source and many presenting as ordinary browsers.
Content delivery networks behave differently
A content delivery network, or CDN, sits between a site's origin server and its visitors, and its rules decide what a blocked request looks like. HasData's table for the 2,096-domain enforcement subset shows four distinct postures toward GPTBot.
| CDN | Served OK | Hard-blocked | Challenge | Rate-limited | Strategy |
|---|---|---|---|---|---|
| CloudFront | 78.5% | 14.4% | 0.0% | 0.0% | Permissive |
| Cloudflare | 51.3% | 16.9% | 24.7% | 0.5% | Challenge-based |
| Akamai | 39.1% | 34.5% | 0.0% | 19.5% | Rate-limit plus block |
| Fastly | 27.3% | 61.5% | 0.0% | 0.6% | Hard block |
By these figures Fastly's hard-block rate is about 3.6 times Cloudflare's, although Cloudflare dominates the policy conversation. Per-CDN site counts are not given, so the sample behind each row cannot be judged.
What Cloudflare's September 15 default covers
Cloudflare said on July 1, 2026 that from September 15 it would block training and agent crawlers by default on ad-carrying pages for newly onboarded domains, and moved its pay per crawl model toward paying per citation. The announcement gave bots as 57.4% of HTML traffic in early June 2026.
HasData measured the footprint of that rule by combining two conditions, a site behind Cloudflare's proxy and ad-tech signals. Only 26.0% of the top 10,000 domains sit behind the proxy (22.4% of publishers); adding the ad-tech condition brings coverage to 8.5% of the top web and 13.6% of publishers, and to 8.6% and 13.3% on September 16. The remaining 91.5% of the top web sits outside the rule, since Fastly, CloudFront, Akamai and origin-only setups are unaffected. The figure counts sites that meet the conditions rather than sites that had the default applied, because the July announcement described newly onboarded domains.
The September 15 update changed the plan. Coverage published on September 27 reports that Cloudflare added a Disallow AI Training option, extended Block to mixed-use crawlers such as Googlebot, Applebot and Bingbot, and deprecated Block AI Bots and Managed Robots.txt in favour of Bot Preference Sync. The planned block on Googlebot for sites refusing training was abandoned. Cloudflare's figures there: 17% of its sites enable some form of training block and under 1% block search. Microsoft's support for a robots.txt no-training signal is targeted for early 2027.
A customer email dated September 16, as the same coverage describes it, said settings would migrate over the following week, so HasData's snapshot, taken between 11:50 and 13:46 UTC that day, falls at the start of that window. HasData, citing Cloudflare's engineering post, says legacy users are prompted to review and confirm Bot Preference Sync, published on August 21, before it takes effect, which differs from an automatic migration.
What the September re-run found
The re-run covered the same 10,894 domains, the same 2,096-site live test and the same 600-site residential sample. The share of sites announcing an AI block in robots.txt fell from 15.2% to 13.0%. Among sites behind Cloudflare, the share disallowing GPTBot fell from 17.1% to 9.9%, while for sites outside Cloudflare the figure moved from 18.7% to 18.6%.
| Metric | Top web, July to September | Publishers, July to September |
|---|---|---|
| Behind Cloudflare | 26.0% to 26.3% | 22.4% to 22.3% |
| In scope (Cloudflare plus ads) | 8.5% to 8.6% | 13.6% to 13.3% |
| Blocks any AI crawler | 10.3% to 8.3% | 56.4% to 53.0% |
| Blocks GPTBot | 7.9% to 5.7% | 50.5% to 46.9% |
| GPTBot got a page | 67.8% to 68.6% | 54.2% to 47.2% |
| Browser got a page | 72.3% to 73.9% | 83.8% to 83.1% |
| Paper-only bans | 46.7% to 59.1% (44-45 sites) | 38.9% to 32.8% |
| Silent enforcers (sites) | 37 to 45 | 78 to 133 |
| Publishes llms.txt | 7.9% to 9.3% | 3.2% to 4.5% |
To see what had been removed, HasData pulled archived copies from Common Crawl's July and August crawls for the 165 Cloudflare sites that stopped disallowing GPTBot, plus a control group of 150 that did not. Copies existed for 133 of the 165. In 118 of those, the archived file carried Cloudflare's managed block (107 from the August crawl, 11 from July), and all 118 showed the same three elements: a Content-Signal line reading ai-train=no, a rights reservation citing Article 4 of the EU copyright directive, and a GPTBot ban. Among the sites HasData names are the Smithsonian, Ko-fi, AirAsia, Find a Grave, Vatican News, National Review, Punchbowl News, Bellingcat, Charlie Hebdo and Novaya Gazeta. On September 16, none of the 118 carried it.
A live re-check the following day found an AI-training signal on one of the 101 sites that answered. That site was Patreon, which writes its own Content-Signal line with ai-train=no, scoped to two named crawlers. Sites that wrote their own rules kept them, HasData says, while those that let Cloudflare write them were left with an empty file. The replacement had not filled the gap: Bot Preference Sync appeared in none of the 6,451 robots.txt files HasData collected.
Does the missing line mean the door opened? Mostly not. Of the 21 affected sites that also ran in the live test, 16 still blocked or challenged GPTBot. Only the machine-readable "no", the part a crawler reads before deciding whether to fetch, disappeared.
In the live test, 125 sites started blocking GPTBot between the two runs and 66 stopped; 112 of the 125 are news publishers and only 22 sit behind Cloudflare, so the default explains a small part of the shift, according to HasData. The browser-to-GPTBot gap on publisher sites widened from 29.6 points in July to 35.9 in September, with the browser row barely moving. Publishers moved in both directions at once: fewer wrote a GPTBot ban, but a larger share of the remaining bans are enforced. The page uses the figure 125 twice - for sites that began blocking GPTBot and, separately, for Cloudflare sites that served robots.txt in July and refused it in September - and does not say whether the two overlap.
Google AI Mode citations and the Google-Extended question
Ten queries run against Google AI Mode in July across a general news mix returned 82 unique cited domains. Of these, 52 appear in the sample, and 27 of the 52 (51.9%) disallow at least one AI crawler in robots.txt, compared with 15% across the whole sample. Among the most-cited blockers, nbcnews.com drew 3 citations with 8 AI bots blocked, nytimes.com, cnn.com, cnbc.com, theguardian.com, linkedin.com and finance.yahoo.com 2 each, and thehill.com, bbc.co.uk and apnews.com 1 each. The test was not repeated in September, and ten queries is a narrow base for conclusions about how Google sources answers.
HasData's explanation is that Google-Extended is a separate directive from GPTBot, so blocking OpenAI's crawler leaves Google's own index available to AI Overviews and AI Mode. The page adds that Google-Extended governs grounding for both. Google's documentation describes the control as covering Gemini training and grounding in Gemini Apps and Vertex AI, not AI Overviews or AI Mode, so the sources differ. That documentation, published on July 20, 2026, also describes a separate Search Console toggle that works at domain level. It followed a conduct requirement imposed by the UK Competition and Markets Authority on June 3, whose main obligations take legal effect on December 3, 2026, with page-level controls due on March 3, 2027.
Citations that persist despite blocking are not new: a BuzzStream study of 4 million citations across 3,600 prompts found that 92.3% of top news sites blocking Google-Extended still appeared in citations.
llms.txt
The index also tracks llms.txt, a proposed format that offers language models a curated list of URLs and short descriptions. Adoption across the sample was 7.4% in July and 8.8% in September; among the open web it moved from 7.9% to 9.3%, and among publishers from 3.2% to 4.5%. The split inverts the robots.txt pattern: publishers block AI crawlers at 5.5 times the open-web rate but publish llms.txt at about 40% of it.
The format has no enforcement side. HasData states that OpenAI, Anthropic, Google and Perplexity read robots.txt, while none has committed to reading llms.txt. Figures published in July showed that Originality.ai counted 36,120 llms.txt instances by May 2026, up 8.8 times year on year, while Ahrefs found 97% of files received zero requests that month.
Why the findings matter to publishers and advertisers
Two ledgers now describe the same site. One is the declared policy in robots.txt. The other is enforcement at the CDN or firewall, which only a live request reveals. In HasData's data the two diverge in both directions: 39.5% of declared bans did not stop a live request, and 5.5% of the subset blocked without declaring anything. Counts of who blocks AI that rely on robots.txt alone therefore measure intent rather than outcome.
Classification adds a second layer of uncertainty. Challenges and CAPTCHAs, 8.2% of HasData's measurements, are folded into either the pass or the block bucket by many public counts, which explains how the same data can yield very different headline figures. The ad-carrying condition in Cloudflare's rule ties the question to revenue, since pages that run advertising are the ones the default targets. Detection tooling is appearing as well: Microsoft added robots.txt violation flags to Clarity's Bot Analytics on June 23, 2026, though it reports violations and does not enforce rules.
Blocking is not free, either. A Rutgers and Wharton paper associated blocking AI crawlers with a loss of roughly 7% of weekly traffic within six weeks, while HasData's citation data and the BuzzStream study indicate that blocking other companies' crawlers does not remove a site from Google's AI answers. A block can carry a measured traffic cost and still leave citations in place.
Several questions remain open. Will the picture change once Cloudflare's migration completes and existing customers confirm Bot Preference Sync? What will Microsoft's robots.txt no-training support look like in early 2027? And how will Google's page-level controls, due March 3, 2027, affect AI Mode citations? HasData's page announces no further re-runs, and the 118-site finding compares archived copies with live checks made while the migration was still under way.
Timeline
- July 1, 2025 - Cloudflare opens pay per crawl in private beta, using HTTP 402 to signal payment requirements to AI crawlers.
- March 19, 2026 - BuzzStream publishes a study of 4 million citations across 3,600 prompts; 92.3% of top news sites blocking Google-Extended still appear in citations.
- June 3, 2026 - The UK Competition and Markets Authority imposes a conduct requirement on Google; Google begins testing a Search Console toggle the same day.
- June 23, 2026 - Microsoft adds robots.txt violation detection to Clarity's Bot Analytics.
- July 1, 2026 - Cloudflare announces default blocking of training and agent crawlers on ad-carrying pages from September 15 for newly onboarded domains.
- July 9 and 10, 2026 - HasData draws its sample from Tranco list 46XLX (dated July 9) and measures the July baseline.
- July 20, 2026 - Google documents a domain-level Search Console control for AI Overviews and AI Mode.
- August 14, 2026 - TollBit data shows about 15% of identified AI page fetchers in Europe reached disallowed URLs.
- August 21, 2026 - Cloudflare publishes Bot Preference Sync.
- September 15, 2026 - Cloudflare's AI crawler update takes effect, deprecating Managed Robots.txt and Block AI Bots.
- September 16, 2026 - HasData re-runs the study between 11:50 and 13:46 UTC; a Cloudflare customer email of the same date says settings will migrate over the following week.
- September 29, 2026 - HasData's AI Crawler Block Index page shows its last update.
- December 3, 2026 - Main obligations of the CMA's conduct requirement take legal effect.
- Early 2027 - Target for Microsoft's support of a robots.txt no-training signal.
- March 3, 2027 - Page-level grounding controls for Google's search generative AI features are due.
Related PPC Land coverage
- Cloudflare drops planned Googlebot block for sites refusing AI training - Covers the September 15 update that deprecated Managed Robots.txt and extended Block to mixed-use crawlers.
- Cloudflare blocks opaque AI crawlers from sites that disallow training - Describes Bot Preference Sync, published August 21, 2026.
- Cloudflare stops charging AI per crawl and starts paying per answer - Reports the July 1, 2026 announcement behind the September 15 default.
- Cloudflare launches pay per crawl to monetize AI content access - Explains the HTTP 402 mechanism introduced in private beta on July 1, 2025.
- 15% of AI page fetchers in Europe reached disallowed URLs, TollBit finds - Gives an independent measure of robots.txt compliance by AI bots.
- Google gives site owners a toggle to exit AI Overviews and AI Mode - Documents the domain-level Search Console control published July 20, 2026.
- UK regulator forces Google to give publishers AI opt-out rights today - Sets out the CMA's June 3, 2026 conduct requirement and its December 3, 2026 and March 3, 2027 dates.
- Blocking AI crawlers doesn't stop citations, new data shows why - Analyses 4 million citations and the persistence of blocked sites in AI answers.
- Blocking AI crawlers cost news publishers 7% of traffic, study finds - Summarises the Rutgers and Wharton estimate of traffic losses after blocking.
- llms.txt adoption rises 8.8x but 97% of files get zero AI requests - Adds adoption and request data for the advisory llms.txt format.
- The user-agent strings every SEO and site owner needs right now - Lists crawler identifiers, including GPTBot and OAI-SearchBot, and how robots.txt applies to them.
- Microsoft Clarity now flags robots.txt violations inside Bot Analytics - Covers the June 23, 2026 addition of violation detection.
- AI agents caught masquerading as humans to bypass website defenses - Describes a Grok query that produced 16 requests from 12 IP addresses.
Summary
- Who: HasData, a web scraping API provider whose co-founder and CTO, Roman Milyushkevich, wrote the index; Cloudflare, whose September 15, 2026 change is the subject of the re-run; and the 10,894 websites in the sample, split between the Tranco top web and news publishers.
- What: An update to the AI Crawler Block Index comparing a July 2026 baseline with a September 16 re-run. The index finds that 39.5% of sites disallowing GPTBot still served it a page, that publishers block AI crawlers 5.5 times as often as the open web, and that Cloudflare's managed AI block was gone from 118 sites' robots.txt files, though 16 of 21 tested affected sites still blocked or challenged the crawler.
- When: The July baseline ran around July 9 and 10, 2026, the re-run took place on September 16, 2026, and the page shows a last update of September 29, 2026.
- Where: Requests came from a US datacenter IP, with a 600-domain residential sample, against sites worldwide. Archived copies from Common Crawl's July and August crawls provided the pre-change robots.txt files.
- Why: Declared policy in robots.txt and actual enforcement at the CDN or firewall layer diverge, and Cloudflare's replacement of Managed Robots.txt with Bot Preference Sync changed what a crawler can read before it fetches a page. The report matters to publishers and advertisers because counts of AI blocking depend on which layer is measured.
Discussion