The United States originates more than half of the world's automated web requests, according to analysis released today by Decodo, the web data collection firm formerly known as Smartproxy. The ranking arrives roughly two months after automated requests passed human ones on the open web for the first time on record.
The figure is 53.5%. No other country comes close. Germany, the second-largest source, accounts for 8.2% of global bot traffic, followed by the Netherlands at 5.6% and Singapore at 5.4%. France and China tie at 3.9% each, India at 3.5%, Ireland at 3.4%, the United Kingdom at 3.3% and Brazil at 3.2%.
Those numbers describe origin, not behaviour. A request tagged to Ashburn or Frankfurt carries the flag of the data centre it left from, not the nationality of whoever, or whatever, initiated it. That distinction runs through the entire dataset and explains most of what looks surprising in it.
A ranking that tracks server capacity
Bot origin follows cloud infrastructure rather than population. Crawlers, agents and proxies run on rented machines, and those machines cluster in a handful of markets with cheap power, dense fibre and permissive hosting economics. India, a country of more than a billion people, contributes 3.5% of global bot traffic. The Netherlands, with a population smaller than that of Mumbai, contributes 5.6%.
The full ranking published by Decodo covers 20 countries.
| Rank | Country | Share of global bot traffic | Bots as % of own traffic |
|---|---|---|---|
| 1 | United States | 53.5% | 43.6% |
| 2 | Germany | 8.2% | 45.0% |
| 3 | Netherlands | 5.6% | 61.3% |
| 4 | Singapore | 5.4% | 73.7% |
| 5 | France | 3.9% | 33.1% |
| 6 | China | 3.9% | 38.4% |
| 7 | India | 3.5% | 14.9% |
| 8 | Ireland | 3.4% | 71.1% |
| 9 | United Kingdom | 3.3% | 21.1% |
| 10 | Brazil | 3.2% | 15.9% |
| 11 | Japan | 3.0% | 17.6% |
| 12 | Russian Federation | 2.7% | 40.5% |
| 13 | Iran | 2.4% | 81.4% |
| 14 | Hong Kong | 2.2% | 55.6% |
| 15 | Indonesia | 1.9% | 18.2% |
| 16 | Canada | 1.9% | 20.8% |
| 17 | Vietnam | 1.7% | 24.8% |
| 18 | South Korea | 1.3% | 21.5% |
| 19 | Finland | 1.1% | 56.8% |
| 20 | Australia | 1.1% | 20.3% |
The second column tells a different story from the first. Iran ranks thirteenth by volume, yet 81.4% of all web traffic attributed to it is automated, the highest proportion in the dataset. Singapore follows at 73.7%, Ireland at 71.1%, the Netherlands at 61.3%, Finland at 56.8% and Hong Kong at 55.6%. By Decodo's count, six of the top 60 traffic-source countries already send more machine requests than human ones.
At the other end sit markets where humans still dominate by a wide margin. India runs 14.9% bot traffic, Mexico 9.2% and the Philippines 10.7%. Readers routed through those networks encounter a measurably different internet from readers routed through northern Virginia.
Sub-national data sharpens the point further. Within the United States, Virginia alone accounts for 27% of national bot traffic, California 9.9% and Oregon 7.7%. England carries 96.3% of UK bot traffic. Hesse leads Germany at 46.1%, and North Holland leads the Netherlands at 55.9%. Each of those regions hosts a concentration of data centres out of proportion to its residential population.
The crossover arrived a year early
Underneath the geography sits a threshold that was crossed in June. Automated systems now generate 57.4% of HTTP requests for web content, against 42.6% from people, a split Cloudflare chief executive Matthew Prince posted on June 3, 2026 using the company's own Radar telemetry. Prince had told an audience at SXSW in March that the crossover would not arrive until 2027. It landed more than a year ahead of that estimate.
PPC Land reported the same Cloudflare Radar measurement from the seven-day window ending June 5, 2026, with bots at 57.4% of HTML traffic and training crawlers alone accounting for 50.6% of the total. Search crawlers made up only 10.7%. That composition matters more than the headline percentage: the majority of machine traffic is not indexing pages for a results page that sends visitors back.
IAB Australia later cited a marginally different Cloudflare figure, putting automated requests at 57.5% of web-page requests as of June 2026 in guidance distributed to members at the end of July. The two numbers describe the same event measured slightly differently.
Growth rates that break analytics assumptions
Decodo attributes the shift to AI systems reading the web at machine speed, drawing its growth figures from HUMAN Security's 2026 State of AI Traffic report rather than its own telemetry.
That report put AI-driven traffic growth at approximately 187% across 2025, roughly eight times the rate at which human traffic grew over the same period. Agentic AI traffic, meaning software acting on a person's behalf rather than collecting training data, grew by approximately 7,851% year over year. PPC Land covered the underlying benchmark in June, noting that just half a percentage point separated benign automation from malicious fraud across the interactions HUMAN processed.
The arithmetic behind the growth is a fan-out problem. A person researching a camera might open three or four pages. An agent completing the same task can issue thousands of requests, because the cost of a fetch is negligible and breadth improves the answer. One instruction becomes a traffic spike.
Decodo splits automated traffic into three groups: training crawlers, which collect data to build and update models and still represent the largest slice of AI traffic; AI agents and fetchers, which pull live pages to answer a prompt; and malicious bots running scraping, credential stuffing and fraud. Roughly a third of bot traffic is classified as malicious in public reporting, leaving about two-thirds doing work that a site operator might reasonably want to permit.
Ownership of the legitimate share is concentrated. Bots operated by OpenAI account for approximately 69% of observed AI-driven traffic by volume, Meta approximately 16% and Anthropic approximately 11%, per the same HUMAN Security data. Three access decisions therefore govern the overwhelming majority of a typical site's AI traffic.
Where the requests land
Sector distribution is where the analysis turns commercially specific. Decodo's own anonymised customer data over the preceding six months places search engines and AI answer engines at roughly 72% of all traffic observed. Retail and eCommerce sites follow as the largest commercial category at approximately 13% of all successful requests.
Everything else trails badly. Travel, airlines and cargo sit at 0.4%, news and media at 0.2%, real estate at 0.2%, jobs and professional data at 0.2%, and finance at 0.1%.
| Industry | Share of total requests |
|---|---|
| Search engines and AI assistants | ~72% |
| Retail and eCommerce | ~13% |
| Travel, airlines and cargo | ~0.4% |
| News and media | ~0.2% |
| Real estate | ~0.2% |
| Jobs and professional data | ~0.2% |
| Finance | ~0.1% |
The pattern reflects volatility. Prices, inventory and listings change hourly, so machines return to them constantly. Static content gets fetched once and cached.
Retail's weight in training crawler activity is heavier still, at 62.5% of that category. Three verticals absorbed more than 95% of AI-driven traffic during 2025: retail and eCommerce, streaming and media, and travel and hospitality.
Route-level data explains why merchants notice. During 2025, 77% of agentic AI activity reached product and search pages, 8.8% account pages, 5% authentication flows and 2.3% checkout. HUMAN Security's May 2026 benchmark, covered by PPC Land in June, found a near-identical distribution, with 76.4% of agentic activity concentrated in product and search routes and checkout at 2.4%, while overall agentic traffic dipped 4.3% month over month and blocking rates climbed toward 9%.
Vaidotas Juknys, chief executive at Decodo, framed the measurement problem in the company's published analysis: "We spent thirty years building the web for people who click. Now, most of the traffic doesn't click, it queries. That's the internet doing exactly what we asked AI to do."
Blocking, charging and the identity problem
The operational question the data raises is not whether to admit automation but how to sort it. Juknys argued against blanket filtering: "The instinct to block every bot is understandable, but not every automated visitor is malicious. As AI assistants become another way consumers discover products and services, businesses need to distinguish between harmful bots and legitimate AI systems. The agents crawling websites today may become how customers discover brands tomorrow."
Decodo sells proxy infrastructure and scraping APIs, which places the firm on the collection side of that argument. The position is nonetheless consistent with what platform operators have been building.
Cloudflare moved toward a pay-to-crawl model that returns an HTTP 402 response carrying a price and settles payment before serving the page, a system first opened in private beta in July 2025. Decodo reports that Cloudflare customers now issue more than one billion 402 responses a day. On July 1, 2026, the company shifted away from charging per crawl toward compensating publishers per answer, on the reasoning that more than half of good-bot crawls re-fetch pages that never changed.
The economics rest on a crawl-to-referral ratio: pages taken against visits returned. Google fetches roughly five pages per referral. Some AI crawlers pull thousands. Cloudflare's own attribution data, published on July 1, 2026, spanned ratios from 118 crawls per referral to nearly 50,000.
Charging presupposes identification, and identification is contested. In August 2025, Cloudflare accused Perplexity of stealth crawling, fetching pages while disguising its bot as a standard browser. PPC Land subsequently documented 20 to 25 million daily requests from the declared crawler alongside 3 to 6 million from an undeclared one, a dispute that escalated into Reddit's federal lawsuit in October 2025. A crawler that hides its name can be neither permitted nor invoiced.
Web Bot Auth, a companion standard using cryptographic signatures, exists to close that gap by making user-agent strings unforgeable. Operators have also begun separating their fleets by purpose: Anthropic runs ClaudeBot for training and Claude-SearchBot for live answers, mirroring OpenAI's split between GPTBot and OAI-SearchBot. That separation lets a site admit the crawler that drives discovery while refusing the one that only extracts.
Publishers are already acting on the distinction. This month, the independent site PatronView blocked Amazon's search crawler after it read 117,000 pages a day, while continuing to allow Googlebot and Bingbot.
Methodology and its limits
The analysis draws on Cloudflare Radar together with Decodo's internal data to examine the split between automated and human requests and the geographic distribution of bot traffic. AI traffic growth figures come from HUMAN Security's 2026 State of AI Traffic report, released on April 9, 2026.
One caveat is stated explicitly by Decodo and matters for anyone reading the regional splits. The US state ranking represents each location's share of bot traffic originating from the United States. It does not represent the percentage of each state's total internet traffic that is automated. Virginia's 27% figure describes concentration of origin, not saturation of a local audience.
Two further limits apply. The sector percentages derive from Decodo's own anonymised customer base, which is composed of organisations buying data collection infrastructure, so the distribution reflects what those customers target rather than the web as a whole. And the underlying blog analysis carries a last-updated date of July 8, 2026, meaning the traffic figures predate the pitch by roughly a month.
Why this matters for the marketing community
For anyone measuring a website, the immediate consequence is arithmetic. A session count that mixes thousands of agent fetches with a handful of human visits misstates demand in both directions, and the error compounds through conversion rates, cost-per-visit calculations and content performance reviews. Automated traffic is now the majority of requests, which means the default assumption behind a decade of analytics tooling no longer holds.
Measurement vendors have been closing the gap in stages. Microsoft Clarity added a bot activity dashboard in January 2026 and extended it in June with robots.txt violation detection showing which crawlers ignore stated access rules. HUMAN Security took its agentic visibility tooling to marketing and commerce teams in April 2026, delivered natively inside Adobe Experience Platform.
For retailers, the exposure is sharper still. If 77% of agentic activity lands on product and search pages, and retail absorbs 62.5% of training crawler traffic, then a bot-blocking rule written for fraud prevention also determines whether a catalogue appears inside AI shopping answers. Botify data reported in March 2026 found that AI bots crawl retail sites 198 times more per visit than Google, a ratio that makes the infrastructure cost of admitting them real rather than theoretical. Kinsta's June analysis of 10 billion requests recorded a single crawler hitting WordPress cart URLs 3.75 million times in one day.
The counter-cost is equally documented. Research from Rutgers Business School and The Wharton School found news publishers who blocked AI crawlers through robots.txt lost 23.1% of monthly visits without achieving proportional protection, because the protocol remains voluntary.
There is also a date on the calendar. Cloudflare's default blocking of Training and Agent category crawlers on advertising-carrying pages takes effect for domains newly joining its network on September 15, 2026, with Search category crawlers still permitted by default. That policy, detailed alongside the company's brand visibility scoring released on August 6, 2026, converts an abstract debate about crawler access into a configuration setting with a deadline attached.
Verification is moving to the edge as well. Fastly integrated Experian's agent checks in late July, as automated traffic was measured at 53% of the web by Imperva, resolving agent identity and delegated authority before requests reach origin servers. Traffic sorted at that layer arrives in analytics and retail media systems already classified, which is a measurement change presented as a security change.
Adoption of the commercial layer remains thin against all this infrastructure. An Originality.ai study published in May 2026 found only 26 public websites had implemented the Universal Commerce Protocol out of more than three million scanned. The plumbing has multiplied faster than the transactions running through it.
What the Decodo ranking adds is a map of where the machines physically sit. Half the world's automated requests leave one country, and a quarter of that country's share leaves one state. For a marketer reading a geographic traffic report, that concentration is the difference between a genuine audience signal and a data centre.
Timeline
- July 2, 2025 - Cloudflare opens pay per crawl in private beta, using HTTP 402 responses to price AI crawler access
- August 2025 - Cloudflare accuses Perplexity of stealth crawling, fetching pages while disguising its bot identity
- September 9, 2025 - Decodo publishes scraping research showing TikTok as the most scraped website of 2025 with 321% traffic growth
- October 24, 2025 - Perplexity denies training models as Cloudflare documents 20 to 25 million daily declared requests alongside 3 to 6 million undeclared
- During 2025 - Agentic AI traffic grows approximately 7,851% year over year and overall AI-driven traffic approximately 187%
- January 21, 2026 - Microsoft Clarity launches its bot activity dashboard
- March 2026 - Matthew Prince tells an SXSW audience the bot-human crossover will not arrive until 2027
- March 7, 2026 - Botify analysis finds AI bots crawl retail sites 198 times more per visit than Google
- April 9, 2026 - HUMAN Security releases the 2026 State of AI Traffic report
- April 21, 2026 - HUMAN Security extends agentic visibility to marketing and commerce teams
- June 3, 2026 - Matthew Prince posts Cloudflare Radar data showing automated systems at 57.4% of HTTP requests for web content
- June 4, 2026 - HUMAN Security's May benchmark records agentic traffic down 4.3% month over month with blocking rates near 9%
- June 5, 2026 - Cloudflare Radar's seven-day window puts bots at 57.4% of HTML traffic against 42.6% human
- June 18, 2026 - Kinsta reports 3.75 million daily bot requests to WordPress cart URLs across 10 billion requests analysed
- June 23, 2026 - Microsoft Clarity adds robots.txt violation detection to Bot Analytics
- July 1, 2026 - Cloudflare moves from charging per crawl to paying per answer and publishes crawl-to-referral ratios spanning 118 to nearly 50,000
- July 8, 2026 - Decodo's bot traffic analysis carries its last-updated date
- July 30, 2026 - IAB Australia publishes crawler guidance citing automated requests at 57.5% of web-page requests
- August 6, 2026 - Cloudflare releases its brand visibility scoring and confirms the September 15 default blocking date
- August 11, 2026 - Decodo releases the country ranking placing the United States at 53.5% of global bot traffic
- September 15, 2026 - Scheduled start of Cloudflare default blocking for Training and Agent crawlers on ad-bearing pages for newly joining domains
Related PPC Land coverage
- Bots now outnumber humans on the web and most aren't here to search - Reports the Cloudflare Radar measurement for the week ending June 5, 2026, with training crawlers alone at 50.6% of total traffic.
- Bots overtake humans - Covers IAB Australia's crawler guidance and decision matrix, including DataDome figures on AI-agent request growth.
- AI agents are now buying things and fraud looks identical - Details HUMAN Security's 7,851% agentic growth figure and the narrow signal gap between agents and fraud.
- AI agent traffic is up 8x and HUMAN Security now tells marketers why - Documents the extension of agentic visibility tooling into marketing and commerce workflows.
- AI agent traffic dips in May but blocking rates keep climbing - Breaks down agentic route distribution across product, search and checkout endpoints.
- AI bots crawl retail sites 198x more than Google, new report warns - Quantifies the crawl-to-referral imbalance facing retail and eCommerce operators.
- AI bots hammered WordPress cart pages 3.75M times in a day, Kinsta data shows - Measures the infrastructure cost of crawler loops at the application layer.
- Cloudflare stops charging AI per crawl and starts paying per answer - Explains the July 2026 shift in crawler compensation and the waste finding behind it.
- Cloudflare exposes AI crawlers hitting sites 50000 times per visitor - Sets out the attribution dashboard and per-operator crawl-to-referral reporting.
- Cloudflare scores brand visibility inside Claude and GPT answers - Covers the August 2026 dashboard release and the September 15 default blocking deadline.
- Perplexity denies training AI models as Cloudflare documents stealth crawlers - Chronicles the crawler identity dispute that made verification a technical requirement.
- Microsoft Clarity exposes AI bot traffic with new visibility dashboard - Describes the first mainstream analytics view into per-operator bot activity.
- Microsoft Clarity now flags robots.txt violations inside Bot Analytics - Adds compliance measurement to crawler volume reporting.
- PatronView blocks Amazon's AI crawler after 117,000 daily page reads - Records a publisher-side blocking decision and the analytics undercount behind it.
- Fastly gains Experian agent checks as 53% of web traffic turns automated - Tracks agent verification moving to the content delivery layer.
- All About Berlin loses 75% of traffic as zero-click searches hit 68% - Presents the publisher-side research on what blocking crawlers costs in visits.
- TikTok leads data collection surge as AI training reshapes scraping landscape - Covers Decodo's earlier scraping research and its methodology base.
- Cloudflare launches pay per crawl to monetize AI content access - Documents the original HTTP 402 crawl payment mechanism and Web Bot Auth requirements.
- Bot traffic costing advertisers billions as fraud detection fails, investigation reveals - Establishes the earlier baseline of at least 40% non-human web traffic and verification failures.
Summary
Who: Decodo, the Vilnius-based web data collection infrastructure provider formerly known as Smartproxy, with chief executive Vaidotas Juknys commenting on the findings. The analysis incorporates Cloudflare Radar telemetry and growth figures from HUMAN Security's 2026 State of AI Traffic report.
What: A country-level ranking of global bot traffic placing the United States at 53.5% of worldwide automated requests, Germany at 8.2% and the Netherlands at 5.6%, alongside a second measure of bots as a share of each country's own traffic, where Iran leads at 81.4%, Singapore at 73.7% and Ireland at 71.1%. The dataset also records bots at 57.4% of global web requests against 42.6% from humans, agentic AI traffic growth of approximately 7,851% during 2025, overall AI-driven traffic growth of approximately 187%, and sector concentration placing search engines and AI assistants at roughly 72% of observed requests and retail at approximately 13%.
When: The ranking was released today, August 11, 2026. The underlying analysis carries a last-updated date of July 8, 2026. The bot-human crossover was posted by Cloudflare on June 3, 2026, and the HUMAN Security report it draws on was published on April 9, 2026.
Where: The measurements cover global web traffic observed through Cloudflare Radar and Decodo's own anonymised customer base, with regional breakdowns for the United States, the United Kingdom, Germany and the Netherlands. Virginia accounts for 27% of US bot traffic, England for 96.3% of UK bot traffic, Hesse for 46.1% of German bot traffic and North Holland for 55.9% of Dutch bot traffic.
Why: Automated requests now form the majority of web traffic, which breaks the assumption behind session-based analytics and forces site operators to sort legitimate AI systems from malicious automation rather than filtering both. The commercial stakes concentrate in retail, where 77% of agentic activity reaches product and search pages and 62.5% of training crawler activity targets retail sites, meaning a crawler-blocking rule written for fraud prevention also governs whether a catalogue appears inside AI-generated shopping answers. A Cloudflare policy change on September 15, 2026 attaches a deadline to that decision for newly joining domains.
Discussion