The crawl-to-referral ratio counts how many pages an artificial intelligence platform's automated software collects from a website for every human visitor that platform later sends to it. A ratio of 5 to 1 means five pages taken for one reader delivered. A ratio of 38,000 to 1 means thirty-eight thousand pages taken for the same single reader. The metric exists because the bargain that funded the commercial web for two decades, page access in exchange for traffic, became difficult to observe once assistants began answering questions instead of listing links, and publishers wanted a number rather than a suspicion.
A crawl is a single HTTP request from an automated client for a page. A referral is a visit arriving with evidence of where it came from. Divide the first by the second, normalise to one, and the result reads as a rate of extraction.
How the number is calculated
Cloudflare publishes the reference implementation. Its Radar methodology divides the total number of requests from user agents associated with a given search or AI platform where the response carried a content type of text/html, by the total number of requests for HTML content where the Referer header contained a hostname belonging to that same platform. The result is always normalised to a single referral request.
Three design choices inside that formula shape everything downstream. HTML pages only, on the reasoning that text is what the crawlers want and images or scripts would inflate the numerator with material nobody is arguing about. All of an operator's user agents are aggregated into one platform figure, so a training crawler and a live fetcher acting for a user are counted together under a single company name. And referral traffic originating in Google's autonomous system 15169 is excluded outright, because Chrome speculation rules prefetch pages before anyone has decided to read them, which would count a machine as a person.
The denominator depends entirely on the Referer header, the field a browser sends to name the page a click came from. Nothing obliges a client to send it.
Two measurements sharing one name
The industry now runs at least two versions of the metric, and they are not interchangeable. Cloudflare Radar produces the network-wide, per-operator figures the trade press quotes. A separate product, the Attribution Business Insights dashboard opened to Bot Management customers on July 1, 2026, produces per-site figures and tracks referrals through UTM parameters rather than the Referer header, viewable across 24 hour, 7 day and 30 day windows and broken out by operator alongside a classification of each crawler as Training, Search or Agent.
Microsoft ships a third variant. Its Clarity analytics product carries an AI Scrape-to-Referral Ratio card with an operator-level ranking of which sources send traffic back and which show heavy scraping with little return. Same concept, different denominator, different vantage point. A figure from one cannot be compared with a figure from another.
Origin and evolution
The framing arrived before the measurement. Cloudflare chief executive Matthew Prince described the deterioration on CNBC on May 21, 2025, saying Google had sent one visitor for every two pages crawled a decade earlier, one for every six by late 2024, and one for every fifteen by that spring, with OpenAI moving from roughly 250 to 1 to about 1,500 to 1 over six months.
The metric itself launched on July 1, 2025, the date Cloudflare branded Content Independence Day, as a new panel on Radar's AI Insights page. Cloudflare's own methodology post reported ratios for June 19 to 26, 2025 running from Anthropic at 70,900 to 1 down to Mistral at 0.1 to 1, the second meaning ten referrals for every crawl.
A fuller dataset followed on August 29, 2025. Cloudflare's crawl-to-click gap analysis tracked seven months of monthly ratios. Anthropic fell from 286,930 to 1 in January to 38,065 to 1 in July, a decline of 86.7% that the company linked to Claude gaining web search with clickable citations. OpenAI moved from 1,217 to 1,091, down 10.4%. Perplexity went the other way, from 54.6 to 194.8, a 256.7% increase. Microsoft held between 38.5 and 45.1. Google ranged from 3.8 in January to 22.5 in April before settling at 5.4. ByteDance collapsed from 18 to 0.9, and DuckDuckGo sat at 0.2. PPC Land documented the per-operator spread at the time.
The figures Prince gave verbally and the figures Radar computed do not reconcile cleanly. Google appeared as 15 to 1 in May 2025, as approximately 14 to 1 in August, and as 5.4 in the July Radar table for the same period. The company has never published a bridge between the two series.
What the spread actually measures
The four-order-of-magnitude gap between Google and Anthropic in mid-2025 was not measurement noise. It tracked crawl purpose. Over the twelve months to August 2025, training accounted for roughly 80% of AI crawling, search for 18% and live user actions for 2%. A training crawler has no referral mechanism at all, so its ratio is bounded only by how much it fetches. A search crawler feeds a surface that can produce a link.
That composition has since moved further in one direction. By early June 2026, automated systems generated 57.4% of requests for HTML content across Cloudflare's network against 42.6% from people, with training crawlers alone at 50.6% and search crawlers at 10.7%. Google fetches roughly five pages per referral; some AI operators pull thousands, according to figures cited in PPC Land's analysis of global bot traffic composition. The July 1, 2026 dashboard put the observed range at 118 crawls per referral at the low end to nearly 50,000 at the high end.
Why the number carries commercial weight
The ratio moved quickly from a chart to an input in pricing and access decisions. Cloudflare replaced per-crawl charging with payments tied to citations on July 1, 2026, and set September 15, 2026 as the date from which Training and Agent crawlers are blocked by default on advertising-carrying pages for domains newly joining its network. Time began routing AI operators through a whitelist, part of a broader shift PPC Land examined when Google dismissed a rival crawler standard in the same week.
Individual publishers use it as a blocking trigger. PatronView, a donor research database, measured Claude-SearchBot at 35,000 to 1 on its own logs and blocked it, then blocked Amazon's crawler after recording about 117,000 requests a day.
Blocking is not free. Research from Rutgers Business School and The Wharton School, revised in April 2026, found publishers who blocked large language model crawlers lost roughly 7% of weekly traffic within six weeks, visible in human browsing panel data. An earlier version of the same work put the figure at 23% of monthly visits for large publishers.
Limitations and disputes
The denominator is the weakest part. Cloudflare states plainly that traffic referred by Claude's native app carries no Referer header, that the same is believed to hold for other native apps, and that the calculations may therefore overstate the ratios by an unknown amount. Links marked with the noreferrer attribute and in-app browsers strip the signal too. A published figure of N to 1 is therefore an upper bound on the real ratio. PPC Land made the same point about Microsoft's card: an undercounted denominator makes the exchange read worse than it is.
The numerator has a mirror problem: JavaScript-based analytics miss requests that reach the server, so crawl counts drawn from tag-based tools understate volume.
A third objection is definitional. Scoring a training crawler on referrals judges it against a function it was never built to perform. Referral clicks are also not the only value a publisher receives; the Rutgers and Wharton work attributed the traffic loss from blocking mainly to reduced brand exposure inside AI answers rather than to forgone referral clicks.
Attribution itself is contested. Cloudflare accused Perplexity of stealth crawling in August 2025, documenting an undeclared crawler using a generic browser user agent alongside the declared one. A ratio computed per operator only works if requests can be assigned to operators. TollBit measurement reported in August 2026 found 15% of AI page fetchers in Europe reaching URLs that had been disallowed.
Finally, the numbers are volatile and vantage-dependent. Anthropic's ratio moved by an order of magnitude inside seven months. Cloudflare sits in front of roughly a fifth of websites, and a single site's figure can sit far from the network average.
Adjacent terms
Crawl budget describes how much crawling a search engine will do on a site. It concerns the supply side of the same requests and contains no referral term.
AI referral traffic is an absolute count of sessions arriving from assistant surfaces. It is the denominator of the ratio in isolation, and it can rise while the ratio worsens.
Scrape-to-referral ratio is Microsoft's name for the same construct inside Clarity, computed on its own tracking rather than on network telemetry.
Answer-engine visibility metrics count appearances and citations inside generated answers. They measure the thing a crawl-to-referral ratio cannot see, which is whether content was used at all when no click followed.
Recent developments
Cloudflare has continued attaching products to the metric. It opened a pilot sharing network freshness signals with OpenAI in July 2026, having earlier tracked that operator at 1,091 crawls per referral, and has argued that more than half of crawl traffic from bots it considers legitimate re-fetches pages unchanged since the previous visit, which makes part of every numerator waste rather than consumption. Microsoft's Clarity card, shipped in August 2026, moved the metric out of a network dashboard and into a general analytics product.
Timeline
- May 21, 2025: Matthew Prince states on CNBC that Google's crawl-to-visitor ratio has moved from 2 to 1 a decade earlier to 15 to 1, with OpenAI at roughly 1,500 to 1.
- July 1, 2025: Cloudflare launches the crawl-to-refer ratio on Radar's AI Insights page, reporting a June 19 to 26 range from Anthropic at 70,900 to 1 to Mistral at 0.1 to 1.
- August 28, 2025: Cloudflare adds AI crawler traffic breakdowns by purpose and industry to Radar.
- August 29, 2025: The crawl-to-click gap post publishes seven months of per-operator ratios, with Anthropic at 38,065 to 1 and OpenAI at 1,091 to 1 in July 2025.
- January 2026: Cloudflare's year in review records AI bots at 4.2% of HTML requests across its network during 2025.
- April 21, 2026: Rutgers and Wharton researchers revise their estimate of the traffic cost of blocking AI crawlers to about 7% of weekly visits.
- June 3, 2026: Cloudflare Radar telemetry puts automated systems at 57.4% of HTTP requests for web content against 42.6% from people.
- July 1, 2026: Cloudflare opens the Attribution Business Insights dashboard, recording ratios from 118 to nearly 50,000, and replaces per-crawl charging with payment tied to citations.
- July 30, 2026: IAB Australia publishes crawler guidance citing automated requests at 57.5% of web-page requests as of June 2026.
- August 2026: Microsoft adds an AI Scrape-to-Referral Ratio card to Clarity.
- September 15, 2026: Scheduled start of Cloudflare's default blocking of Training and Agent crawlers on advertising-carrying pages for newly onboarded domains.
Related PPC Land coverage
- Cloudflare exposes AI crawlers hitting sites 50000 times per visitor - The Attribution Business Insights dashboard, its UTM-based referral counting and the Training, Search and Agent classification.
- AI crawling data reveals massive imbalance in training versus referral patterns - The January to July 2025 per-operator ratios and the growth of training-purpose crawling.
- Cloudflare CEO warns zero-click searches threaten content creator revenues - Matthew Prince's May 2025 CNBC figures, the first widely circulated version of the ratio.
- Cloudflare expands 402 payment protocol for AI crawler communication - The August 2025 AI Crawl Control expansion and the approximately 14 to 1 search baseline cited alongside it.
- Microsoft Clarity card ranks which AI operators scrape most and refer least - Microsoft's scrape-to-referral card and the denominator problem that inflates any such ratio.
- PatronView blocks Amazon's AI crawler after 117,000 daily page reads - A single publisher's server-log ratios and the firewall decisions that followed them.
- US sends 53.5% of global bot traffic, Decodo analysis finds - Traffic composition figures and the contrast between Google's crawl rate and AI operators.
- Cloudflare gives OpenAI network signals covering 20% of the web - The July 2026 freshness-signal pilot and OpenAI's tracked ratio.
- Google's crawler math turns against it as the open web pushes back - Publisher whitelisting through TollBit and Google's dismissal of rival crawler directives.
- Cloudflare stops charging AI per crawl and starts paying per answer - The move from per-crawl pricing to citation-based compensation.
- Cloudflare ties AI payouts to citations as 50% of crawls waste - The September 15, 2026 default-blocking date and the wasted-crawl finding behind it.
- Blocking AI crawlers cost news publishers 7% of traffic, study finds - The revised Rutgers and Wharton estimate and its three-source methodology.
- Blocking AI crawlers backfired: news publishers lost 23% of traffic - The earlier version of the same study, with a materially larger effect size.
- Perplexity denies training AI models as Cloudflare documents stealth crawlers - Undeclared crawling and the attribution problem underneath per-operator ratios.
- Cloudflare blocks opaque AI crawlers from sites that disallow training - Traffic composition by crawl purpose as of early June 2026.
- 15% of AI page fetchers in Europe reached disallowed URLs, TollBit finds - Measurement of the gap between stated access preferences and observed crawler behaviour.
- Cloudflare cuts AI token costs by 80% with markdown conversion - Crawl efficiency work and the ratios cited as its justification.
- Explaining ClaudeBot - The Anthropic training crawler behind the highest ratios in the published series.
- Explaining GPTBot - OpenAI's training crawler and its share of AI crawling traffic.
- Explaining Googlebot - The search crawler that supplies the comparison baseline in every version of the metric.
Summary
Who: Cloudflare originated and publishes the metric, with chief executive Matthew Prince as its most prominent advocate. Microsoft ships a variant inside Clarity. Publishers, media owners and the buyers trading against their inventory are the intended audience, and AI operators including OpenAI, Anthropic, Perplexity, Google and Microsoft are the subjects.
What: A ratio dividing the number of HTML page requests made by an operator's crawlers by the number of visits arriving with that operator named as the referrer, normalised to one referral.
When: Framed publicly in May 2025, launched as a Radar metric on July 1, 2025, expanded into a per-operator time series on August 29, 2025, and turned into a per-site product on July 1, 2026.
Where: Computed across Cloudflare's network, which sits in front of roughly a fifth of websites, and separately inside site-level analytics products from Cloudflare and Microsoft.
Why: Search crawling implied a return: pages read, visitors sent. Answer engines break that loop by satisfying the question in place. The ratio puts a figure on the resulting asymmetry, which is why it now appears in pricing models, blocking policies and licensing arguments, and why its measurement flaws matter as much as its headline values.
Discussion