A small philanthropy database in Austin, Texas, has published a detailed account of a year spent fighting automated traffic, reporting that bots accounted for more than 99% of requests hitting its servers and that one AI search crawler read its pages 35,000 times for every human visitor it sent back.
Nick Gray, founder of PatronView, a research database covering American museum and cultural institution donors, published the findings on August 7, 2026, in a post titled "99% of My Website Traffic Is Bots." The piece walks through a year of server logs, firewall changes, and crawler behavior across a 1.5 million page site built from IRS 990 forms, donor walls, and annual reports.
Gray posted a summary on X the same day, writing that he was "shocked to see that 99% of my traffic is from bots" and that Anthropic's crawler "scrapes my site 35,000 times for every 1 visitor they send." The post drew replies from other site operators describing similar experiences, including Arvid Kahl of Podscan, who said most of his operational work now involves preventing his servers from being overwhelmed by bot traffic.
A server log tells a different story than an analytics dashboard
The core measurement separates what visitor analytics tools report from what a server actually answers. In the week Gray published the piece, PatronView's server answered 2.5 million requests and served 1.28 million full pages, while the site's self-hosted Plausible analytics recorded 5,977 human pageviews, a ratio of roughly 214 non-human page loads for every page load counted as human.
Gray attributes the gap to how JavaScript-based analytics tools work. Plausible, like Fathom or Google Analytics, only counts visitors whose browsers execute JavaScript, and almost no bot does, so a site owner checking only that kind of dashboard has no visibility into what is actually striking the server.
The site itself is built by scraping public philanthropic disclosure documents, a detail Gray addresses directly, distinguishing his own periodic collection of source documents from the thousands of automated requests his own site receives daily.
National traffic spikes preceded the crawler findings
Before Gray isolated individual AI crawlers, the site experienced traffic events tied to geography rather than a single company. In November 2025, roughly four thousand sessions appeared over a few days, each visiting exactly one page with a 99% bounce rate and no referrer, concentrated on fund pages that ordinarily draw only about 10% of real visitor traffic.
The larger spike came on April 22, 2026, when the site recorded 3.6 million requests in a single day, originating from 361,844 unique IP addresses, the large majority located in China. Cloudflare's Managed Challenge system, described in the post as an invisible CAPTCHA, absorbed 1.18 million of those requests within the first ten hours, though a notable share was passing the challenge rather than failing it.
Gray responded by blocking the entire country of China at the network edge on April 23, then extended the same block to Vietnam and Singapore after similar patterns appeared from those countries. He supports the decision with his own audience data: PatronView's genuine search traffic runs 95.9% United States and 1.3% Canada, for a database covering American donors published in English.
Replies to Gray's public post about the incident, cited in his write-up, suggest the pattern is widespread among independent site operators. Matt Paulson of MarketBeat recommended adding Russia to any block list. Jack Ellis of Fathom Analytics reported that customers had seen a comparable wave of Chinese-origin spam traffic over roughly six months before it dropped off entirely. Jeremy Brandt described his own country block list as "extensive." Rodrigo Rocco flagged a harder problem that recurs later in Gray's findings: traffic increasingly arrives through thousands of residential IP addresses making single calls each, defeating geography-based and volume-based blocking simultaneously.
Measuring the crawl-to-referral ratio directly
The most specific figures in Gray's post concern AI crawlers built to power search and answer products rather than general-purpose training crawlers. Cloudflare has previously stated, in figures Gray cites, that Anthropic's crawlers run at roughly 3,000 pages crawled for every one visitor referred. Gray's own measurement, taken in June 2026 using Cloudflare's AI crawler dashboard, put the figure for his site at 35,000 to 1.
The underlying numbers came from two separate user agents Anthropic operates. Claude-SearchBot, the crawler Anthropic uses to build search results, requested 420,680 pages from PatronView over one week. In that same week, Claude-User, the distinct user agent Anthropic sends when someone using Claude asks it to fetch a specific page on a person's behalf, delivered 12 human visitors to the site. Bandwidth told the same story from a different angle: the site served 4.63 gigabytes to the search crawler that week and 175 kilobytes to the humans it referred.
Gray disclosed that PatronView was itself built with Claude Code and that he continues to use Claude Code to refine his Cloudflare security rules, a detail he presents as an acknowledged irony, before describing his decision to block Claude-SearchBot at the firewall level. Following the block, the crawler's request volume fell from roughly 60,000 requests a day to about 25 attempts a day, a pattern he read as evidence that Anthropic's crawler respects a standard HTTP 403 access-denied response rather than attempting to circumvent it.
That measurement produced a broader metric Gray now applies to every crawler hitting the site: pages crawled per human visitor referred. Googlebot measured at 46 crawls per visitor, a ratio Gray describes as earning its keep. Bingbot measured at 406 to 1, worse but still defensible in his framing. The AI crawlers built for training large language models measured far higher, with Claude-SearchBot at 35,000 to 1 before the block.
A second crawler with no referral path at all
While preparing the post, Gray identified Amzn-SearchBot, the crawler Amazon operates to feed its Rufus shopping assistant and Alexa's answer products, as the site's highest-volume crawler at approximately 117,000 requests per day, nearly double what Claude-SearchBot had reached in June before its own block took effect. Because the crawler exists to generate answers inside Amazon's own products rather than to send traffic elsewhere, Gray wrote that it would never refer visitors back and that he doubted it would even provide attribution. He blocked it two days before publishing the post, describing the change as taking two minutes to implement.
Gray also credits Cloudflare's managed AI Crawl Control feature with blocking declared training crawlers, including GPTBot, ClaudeBot, CCBot, and Bytespider, before his own custom rules run. The same dashboard showed Bingbot requesting 158,610 pages against 680 referred visitors, a ratio Gray characterized as acceptable given that his Bing referrals continue to grow.
A CAPTCHA feature that cost more than the site it protected
Separately from the crawler-blocking work, Gray described discovering that Cloudflare's JavaScript Detections feature, which had been injecting a challenge script into every page load throughout the year, was consuming 2,875 milliseconds of load time on a mid-range phone against 278 milliseconds for the site's own JavaScript. He identified the script as the primary reason his mobile Lighthouse performance score sat at 58, noting that the feature's diagnostic output was not readable through his firewall rule interface.
Gray disabled the feature on August 5, 2026, and reported a Lighthouse score of 99 within an hour. Five hours after the change, a scraper operating from Microsoft Azure retrieved 23,000 pages in a single hour, spreading requests across more than 80 IP addresses that each stayed under Gray's existing rate limit, evidence he read as showing the disabled feature had in fact been suppressing that traffic. A separate wave of residential IP addresses began presenting as Chrome versions 118 through 120, browser releases from 2023, following the same disabling.
Residential IP traffic defeats geography and volume-based rules
Gray connected the August pattern to a structurally similar event from November 2025, in which traffic arrived from roughly forty countries in near-identical numbers on the same day, a distribution he characterizes as the signature of a rented residential proxy network in which thousands of ordinary home internet connections are used, by the request, to distribute a single customer's traffic so it appears to originate from real individual users.
By late July 2026, the pattern reappeared, concentrated on American residential internet connections rather than distributed internationally. Unique IP addresses hitting the site climbed from a baseline of roughly 18,000 per day to 124,000 on July 31, 2026, while nothing about the underlying site had changed during the month. Because the traffic originated from ordinary American residential internet service providers rather than from data centers or outside North America, it passed through both of the geography-based and infrastructure-based rules Gray had already built.
His response added two further rules: extending the existing datacenter challenge to cover Azure and other major cloud providers, and introducing what he calls a check against browsers frozen years out of date. Real visitor traffic showed only 0.54% running browser versions old enough to trigger the new rule, with most of that legitimate share isolated to a long-term support release of Firefox that he exempted from the check.
The published rule set and its measured accuracy
Gray's post includes the complete set of nine Cloudflare Web Application Firewall rules currently running on the site, along with the underlying rule-language expressions in an appendix, covering country blocks, named-crawler blocks, a skip rule for Cloudflare-verified bots, continent-level challenges, datacenter-ASN challenges, stale-browser challenges, and a rate limit applied to extensionless page requests. The configuration runs on Cloudflare's Pro plan, which Gray reported costs 25 dollars a month.
To test whether the more aggressive challenge rules were affecting real visitors, Gray published solve-rate data covering a 48-hour window in which Cloudflare issued 106,437 challenges, of which 252 were solved, a rate of 0.24%. Country-level solve rates in the same window ranged from a 99.14% failure rate in India to a 100% failure rate in Iraq, a pattern Gray reads as confirmation the challenged traffic was overwhelmingly automated rather than composed of inconvenienced legitimate visitors.
Gray also disclosed cost figures for running the site. Normal monthly infrastructure spending sits around 90 dollars, running on Cloudflare Workers with a D1 database and edge KV caching, meaning most bot requests hit cached responses at minimal marginal cost. During one period of heavy scraping activity, that bill rose by roughly 500%. Gray stated that on a traditional virtual private server billed by CPU and bandwidth consumption, the traffic volume he documented would represent an existential cost problem, whereas on his current edge-caching architecture it functions as an operational nuisance.
Findings the post frames as unresolved
Gray's concluding section reports that in the 24 hours following implementation of his newest rules, Cloudflare blocked 46,729 requests outright, of which 43,150 originated from Amazon's crawler alone, continuing to strike the block implemented two days earlier. The same period saw 63,969 challenges issued, of which 552 were solved.
He frames the underlying problem as economic rather than technical, pointing to Cloudflare's pay-per-crawl framework, under which crawlers pay a fee per request at the network edge, as a mechanism he would use were it commercially available to him, stating he would be willing to sell Amazon's crawler access to his 3.5 million monthly page views at a fair rate rather than blocking it outright. Absent that kind of market, his stated operating rule is that any crawler that never sends him a visitor gets blocked.
Why this matters for the marketing community
Gray's post is a single site's data, not an industry study, but it lands inside a body of measurement that PPC Land has tracked since mid-2024, and the numbers he reports sit inside ranges prior reporting has already established as directionally consistent across much larger datasets.
The 35,000 to 1 figure for Claude-SearchBot falls within the range Cloudflare itself disclosed when it opened a Bot Management dashboard to customers on July 1, 2026, documenting crawl-to-referral ratios spanning from 118 at the low end to nearly 50,000 at the high end across the crawlers it tracks. Cloudflare's own historical figures for Anthropic specifically moved from 286,930 crawls per referral in January 2025 down to 38,000 by July of that year, according to reporting tied to Anthropic's crawler documentation update published February 25, 2026, which clarified the separate roles of ClaudeBot, Claude-User, and Claude-SearchBot and stated that all three respect robots.txt. Gray's post appears to corroborate that claim directly, reporting that Claude-SearchBot's request volume dropped by roughly 99.96% within days of a firewall block, consistent with a crawler that honors an access-denied response rather than circumventing it.
The Amazon figures fit a similarly documented pattern. A March 2026 report covered by PPC Land found that AI bots crawl retail sites 198 times more per visit than Google, and noted that Amazon's Rufus assistant had by that point generated close to 12 billion dollars in incremental annualized sales, a commercial incentive that helps explain why Amzn-SearchBot's crawl volume on Gray's site nearly doubled Claude-SearchBot's earlier pace.
The residential proxy pattern Gray describes connects to a Samsung security disclosure reported by PPC Land, in which a researcher found proxy code from Bright Data embedded inside a quarter of sampled Tizen smart TV applications, converting consumer devices into exit nodes for the same kind of distributed scraping traffic.
The broader debate over whether blocking AI crawlers helps or harms a publisher remains unsettled in the research PPC Land has tracked. An April 2026 working paper from researchers at Rutgers Business School and The Wharton School found that news publishers who blocked large language model crawlers through robots.txt lost roughly 7% of weekly traffic within six weeks, a finding that stood in some tension with an earlier version of the same research, covered separately, measuring a 23% decline. Gray's post does not directly engage that literature, and his site's traffic composition, a specialist donor database rather than a general news publisher, may not generalize to the outlets that research examined.
Cloudflare's own July 1, 2026 policy shift, moving from charging AI crawlers per individual fetch toward paying publishers based on whether their content was actually used to generate an answer, was tied by the company to internal data showing more than half of crawl traffic from bots it classifies as legitimate re-fetches pages unchanged since a previous visit, a wasted-crawl problem structurally similar to what Gray documents on his own, much smaller site. A companion policy sets a September 15, 2026 default under which Training and Agent category crawlers will be blocked automatically on advertising-carrying pages for domains newly joining Cloudflare's network, while Search category crawlers remain allowed by default, a distinction that maps onto Gray's decision to block Claude-SearchBot and Amzn-SearchBot while continuing to allow Googlebot and Bingbot.
Gray's central methodological point, that visitor-tracking tools built on JavaScript execution systematically undercount traffic actually reaching a server, carries direct implications for anyone relying on tools like Google Analytics or Plausible to assess infrastructure load or the true cost of serving content to non-paying automated visitors.
Timeline
- November 2025 - PatronView records a wave of roughly 4,000 single-page sessions with no referrer, later identified as an early bot pattern.
- January 2025 - Cloudflare data shows Anthropic's crawl-to-referral ratio at 286,930 to 1, later cited in PPC Land coverage of Anthropic's documentation update.
- July 2025 - Anthropic's crawl-to-referral ratio falls to 38,000 to 1, according to the same Cloudflare figures.
- February 25, 2026 - Anthropic clarifies the separate functions of ClaudeBot, Claude-User, and Claude-SearchBot and commits to respecting robots.txt.
- March 2026 - A report finds AI bots crawl retail sites 198 times more per visit than Google, with Amazon's Rufus assistant tied to nearly 12 billion dollars in incremental annualized sales.
- April 21-23, 2026 - PatronView records 3.6 million requests in one day from 361,844 unique IP addresses, mostly in China; Gray blocks China, then Vietnam and Singapore.
- April 26, 2026 - Updated Wharton and Rutgers research finds publishers blocking AI crawlers lost roughly 7% of weekly traffic within six weeks.
- June 2026 - Gray measures Claude-SearchBot's crawl-to-referral ratio on PatronView at 35,000 to 1 and blocks the crawler at the firewall.
- July 1, 2026 - Cloudflare opens a Bot Management dashboard documenting crawl ratios up to 50,000 to 1 and separately announces a shift to per-answer publisher payments, alongside a policy setting a September 15, 2026 default block for Training and Agent crawlers on new domains.
- Late July 2026 - Unique IP addresses hitting PatronView climb from a baseline of roughly 18,000 to 124,000 on July 31, traced to a residential proxy network.
- August 5, 2026 - Gray disables Cloudflare's JavaScript Detections feature, raising his mobile Lighthouse score from 58 to 99 within an hour.
- August 5-6, 2026 - Gray blocks Amazon's Amzn-SearchBot crawler, then measured at approximately 117,000 requests per day.
- August 7, 2026 - Gray publishes the full findings on PatronView and summarizes them on X.
Related PPC Land coverage
- Cloudflare exposes AI crawlers hitting sites 50000 times per visitor - Covers the July 1, 2026 Cloudflare dashboard launch and the wasted-crawl data behind the company's shift to per-answer publisher payments.
- Anthropic clarifies what its three web crawlers do - and how to block them - Details Anthropic's February 2026 documentation separating ClaudeBot, Claude-User, and Claude-SearchBot, and its robots.txt commitments.
- AI bots crawl retail sites 198x more than Google, new report warns - Reports on AI crawler volume against retail sites and the commercial scale of Amazon's Rufus assistant.
- Blocking AI crawlers cost news publishers 7% of traffic, study finds - Covers the Rutgers and Wharton research on the traffic cost of blocking AI crawlers via robots.txt.
- Samsung bans proxy SDKs as a quarter of Tizen apps route strangers' traffic - Documents how residential proxy networks recruit consumer devices as scraping exit nodes.
- Google's crawler math turns against it as the open web pushes back - Analyzes the range of crawl-to-referral ratios Cloudflare's dashboard revealed across AI operators.
Summary
Who: Nick Gray, founder of the philanthropy donor research site PatronView, based in Austin, Texas.
What: Gray published server log data showing that bots accounted for more than 99% of traffic to his 1.5 million page site, including a measured 35,000 to 1 crawl-to-referral ratio for Anthropic's Claude-SearchBot and a 117,000 requests per day rate for Amazon's Amzn-SearchBot, alongside the specific Cloudflare firewall rules he built in response.
When: Gray published the findings on August 7, 2026, describing events across roughly a year, from an initial bot pattern in November 2025 through a firewall change made two days before publication.
Where: The findings concern PatronView's own infrastructure, running on Cloudflare's network, though the traffic sources documented span China, Vietnam, Singapore, and distributed residential internet connections primarily in the United States.
Why: The post provides one of the more granular, independently measured accounts of how AI search crawlers behave on a real production site, offering figures that align with broader industry measurements from Cloudflare and third-party researchers while giving site operators a specific, documented set of firewall rules and their measured effects on legitimate traffic.
Discussion