IAB Australia today released a Bots and Crawler Guidance and Decision Matrix, a document that asks publishers to stop treating every automated visitor as the same problem and instead sort each one by what it actually does with the content it takes. The guidance, unveiled at the trade body's Discovery: AI and Search Summit in Sydney, arrives at a moment the document itself describes starkly: in mid-2026, automated requests overtook human ones in web-page traffic for the first time on record, and the majority of that identified crawler activity now traces back to artificial intelligence systems rather than conventional search engines.
The matrix does not ask publishers to block everything, nor does it ask them to allow everything. Instead, it splits non-human web traffic into five distinct jobs and offers four possible verdicts for each: allow, allow with conditions, require licensing, or block. According to IAB Australia, volume and value have become separate questions, since AI-driven sources often send far less traffic than classic search engines, yet that smaller volume frequently converts at a higher rate.
Five jobs, one decision framework
The guidance divides crawlers into categories the document labels A through E. Category A covers discovery and search indexers such as Googlebot, bingbot, and OAI-SearchBot, described as the librarian that reads a site, catalogues it, and sends readers back. Category B covers AI training crawlers, including GPTBot, ClaudeBot, and CCBot, framed as the student who memorises a text and never returns it. Category C covers live AI agents such as ChatGPT-User, Claude-User, and Google-Agent, fetching a page in real time because a human asked a question. Category D covers operational and advertising infrastructure, including AdsBot-Google and verification fetchers, described as the meter reader that keeps the ad stack functioning. Category E covers unknown or unverified traffic, including spoofed user agents and default library strings such as python-requests.
For each category, the document assigns a publisher trend and a brand trend. Discovery and search crawlers get an allow recommendation across both groups. AI training crawlers get a block-or-license recommendation for publishers, while brands are advised toward a more selective stance, letting product facts through while withholding premium intellectual property. Live AI agents receive a conditional recommendation for publishers, restricted on ad-funded pages, while brands are advised to allow access on product and knowledge pages. Operational infrastructure receives an unconditional allow for both groups, provided the crawler's identity has been verified. Unknown and unverified traffic receives a challenge-or-block recommendation across the board.
According to IAB Australia, the framework draws on the IAB Tech Lab's Bot and Crawler Management Guidance, alongside the Content Monetization Protocol version 1, which the Tech Lab finalised on April 28, 2026. The guidance instructs publishers to decide per crawler group first, then per individual crawler within higher-risk groups, and to document each decision with a rationale and a scheduled review date.
The traffic numbers behind the framework
The document leans on figures the guidance attributes to several named research sources to justify its urgency. Cloudflare data cited in the guidance puts automated requests at 57.5% of web-page requests as of June 2026, describing it as the first such crossover on record. The same Cloudflare figures show roughly 52% of AI crawler requests going toward model training, with only about 2.6% representing real-time, human-triggered fetches. DataDome figures cited in the document show AI-agent requests climbing 45% quarter on quarter, reaching 17.7 billion in the second quarter of 2026.
The guidance also cites DataDome data showing that 80 to 88% of AI referral traffic to sites now originates from ChatGPT, even as the volume of ChatGPT's own crawling activity declined over the same period. That divergence between crawl volume and referral volume runs through the entire document. The guidance states plainly that a person shopping for a pair of shoes might visit five websites, while an agent completing an equivalent task might visit hundreds, and that automated traffic overall is expanding several times faster than human browsing activity, according to HUMAN Security figures the guidance cites for 2026.
Mixed-purpose crawling - a single crawler serving several jobs at once - fell from roughly 49% to 33% of AI requests across the first half of 2026, per the figures cited in the guidance, as operators split their do-everything bots into narrower, single-purpose tokens. Meta's indexing crawler overtook its own training crawler in June 2026, an inversion the document points to as evidence that any allowlist a publisher builds today may already be outdated by the time it gets published.
A note on agentic commerce infrastructure
The guidance also flags Model Context Protocol traffic, describing volume that went from negligible to peaks near 500,000 requests a day, according to DataDome figures cited in the document. IAB Australia frames this as an early signal of agents taking inventory of what they can act on before acting. The guidance cites a Gartner projection that AI agents could intermediate more than 15 trillion US dollars in business-to-business purchasing by 2028, alongside an Adobe figure suggesting AI-referred traffic converts around 42% better than non-AI visits. The document cautions that these figures come from each provider's own network and measurement window and should be read as directional rather than precise.
Training and search are not the same decision
One mechanic recurs throughout the guidance as its central technical point: a single crawler can serve more than one purpose, and blocking a training-only signal does not necessarily remove a site from AI answer features. The document draws a sharp distinction between control tokens such as Google-Extended and Applebot-Extended, which never independently visit a site but are instead robots.txt directives honoured by an operator's main crawler, and the crawler itself. Their scope, according to the guidance, is model training only, meaning they do not remove a publisher's content from answer features built on the underlying search index.
This distinction connects to reporting on Anthropic's own crawler documentation. Anthropic operates three separate bots: ClaudeBot for model training, Claude-User for real-time user queries, and Claude-SearchBot for search result quality, a separation the company clarified in documentation updated around February 20, 2026. Restricting ClaudeBot signals that a site's future material should be excluded from training datasets, while restricting Claude-User prevents the system from retrieving content in response to a live user question, a distinction with different practical consequences for a publisher weighing data rights against audience reach.
The guidance also flags a divergence in how operators treat robots.txt for live, user-triggered fetchers - what the document calls "three ways" of handling the same underlying question. According to the guidance, Google documents that its user-triggered fetchers, including Google-Agent, generally ignore robots.txt because a person requested the page. OpenAI's documentation states that robots.txt may not apply to ChatGPT-User for the same reason. Anthropic, by contrast, states that all of its bots, including Claude-User, will honour robots.txt directives. The guidance's conclusion is unambiguous: for this category, robots.txt is not a reliable control, and publishers should check each operator's current documentation and enforce decisions at the network edge where it matters.
Verifying identity, not just names
A meaningful portion of the guidance addresses what it terms impostor traffic. User-agent strings are voluntary text, the document notes, meaning any scraper can claim to be any operator's crawler. The guidance recommends a three-step verification process: matching the user-agent string against a maintained directory, confirming the source IP address against an operator's published ranges or forward-confirmed reverse DNS, and confirming that behaviour fits the declared purpose, including crawl rate, paths accessed, and robots.txt compliance. A claimed major-crawler user agent arriving from a residential internet service provider or a generic cloud range is, according to the guidance, almost certainly an impostor.
The document also warns publishers against blocking unidentified traffic reflexively, noting that some unrecognised requests originate from a publisher's own procurement stack: content management systems, site search vendors, translation tools, tag management platforms, or data-feed partners. Cross-referencing high-volume unknown traffic against internal procurement and engineering records, the guidance suggests, can prevent a publisher from accidentally severing a service it is already paying for.
The Cloudflare deadline embedded in the guidance
The guidance carries an operational deadline tied to a separate infrastructure provider's policy change. Before September 15, 2026, site owners operating behind Cloudflare or any managed edge or web application firewall are advised to review their AI-crawler defaults. New default settings will block crawlers the guidance classifies as Training and Agent on advertising-carrying pages for new domains and free-tier zones, with multi-purpose crawlers assessed under the most restrictive applicable rule once that date arrives.
This September 15 policy traces to Cloudflare's own July 1, 2026 announcement, covered by PPC Land, which set that date as the start of default blocking for Training and Agent crawlers on ad-bearing pages for domains newly joining Cloudflare's network, while leaving Search crawlers allowed by default under the same rule. The same July 1 announcement introduced a dashboard called Attribution Business Insights, which PPC Land examined in a July 5 analysis, showing crawl-to-referral ratios ranging from 118 crawls per referral at the low end to nearly 50,000 at the high end for some AI operators. Cloudflare also shifted its own compensation model that day, moving away from charging AI companies per individual crawl and toward a structure that pays publishers when their content demonstrably contributes to a generated answer, a change PPC Land reported the same week.
Licensing as a third option, not a new crawler type
The guidance frames licensed access as an emerging alternative to the traditional binary of allowing or blocking traffic outright. Licensing, according to the document, is not a new category of crawler but a new kind of relationship that can apply to any of the five job categories once a commercial agreement exists. The Content Monetization Protocol, finalised by IAB Tech Lab on April 28, 2026, is described in the guidance as the standard communication layer for a licensed AI system and a content owner once that agreement is in place - explicitly not a blocking system, a marketplace, or an economic model in itself, but the plumbing that assumes an access-control strategy already exists.
That framing connects to IAB Tech Lab's own published history with the specification. PPC Land reported that IAB Tech Lab opened CoMP version 1.0 for public comment on March 10, 2026, running through April 9, and that the specification requires AI systems to secure commercial agreements with publishers before crawling content. The companion Bot and Crawler Management Guidance referenced throughout the IAB Australia matrix followed a similar public-comment path: PPC Land reported that the guidance opened for comment on May 27, 2026, running through June 25, after IAB Tech Lab identified bot and crawler management as a gap sitting outside the CoMP specification's technical scope.
Australia's copyright backdrop
The guidance situates its recommendations against a specific legal context. According to the document, Australia has no equivalent of the broad, US-style fair use doctrine, relying instead on narrower fair dealing exceptions, none of which covers AI training under current law. The guidance states that industry consultation has moved toward licensing and compensation models through the Copyright and AI Reference Group. The document's own conclusion from this backdrop is direct: content should be used on negotiated terms, not made available by default, and the guidance explicitly frames this section as general information rather than legal advice.
What publishers are asked to measure
Beyond classification, the guidance sets out a five-step operational sequence. The first step is visibility: an audit of at least 30 days of edge or server logs before any policy change, on the reasoning that owners who block blind risk cutting off crawlers that drive genuine referral traffic. The second step evaluates each crawler on value returned, weighing traffic, content use, bandwidth cost, reputational exposure, and strategic value. The third step applies one of the four verdicts per crawler group, with per-crawler decisions reserved for higher-risk categories. The fourth step calls for documentation: a recorded decision, rationale, any conditions attached, and a review date, with weekly checks for new high-volume unknown traffic, monthly bandwidth reviews, and a full quarterly reassessment. The fifth step asks publishers to publish their policy in three forms: a human-readable page, a machine-readable robots.txt file, and a named licensing contact.
The guidance frames this last step as consequential in itself. A visible, published policy, according to the document, turns a passive block into a potential commercial relationship, since a licensing contact gives an AI operator somewhere to go if it wants access on negotiated terms rather than simply working around a block.
Context for the release
IAB Australia describes the guidance as a resource that will be added to over time, rather than a finished standard. The document itself acknowledges that user-agent tokens change frequently and instructs readers to re-verify identities against the IAB and ABC International Spiders and Bots List and each operator's documentation before implementing any rule based on the current version. The guidance is dated version 1.0, prepared July 22, 2026, and credited to Jonas Jaanimagi, described in the document as IAB Australia's Technology Lead.
The release lands amid a broader run of infrastructure activity PPC Land has tracked through 2026. PPC Land reported that Google published a consolidated web crawling overview on March 3, 2026, and separately documented that Google added Google-Agent to its official crawler list on March 20, 2026, alongside a reference to an experimental web-bot-auth protocol using cryptographic signatures. Cloudflare has pursued a parallel cryptographic approach: PPC Land reported that the company published a registry format for bot and agent authentication on October 30, 2025, building on its Web Bot Auth protocol proposal from May 2025, which uses HTTP Message Signatures with public key cryptography rather than relying on self-reported user-agent strings.
Adoption of complementary voluntary standards has lagged. PPC Land reported that llms.txt adoption grew 8.8 times to nearly 39,000 sites by May 2026, even as a separate Ahrefs analysis of server-log data from 137,000 domains found that 97% of those files received zero AI requests during the same month, with SEO audit tools rather than AI assistants accounting for the largest share of the requests that did occur. That gap between stated preference and enforced access is precisely the problem IAB Australia's guidance frames itself as addressing: robots.txt and similar declarations remain advisory, while network-level enforcement through a content delivery network or web application firewall is not.
Timeline
- August 3, 2024 - Reporting shows over 35% of leading websites blocking AI crawlers as publisher concerns about content scraping mount.
- May 2025 - Cloudflare shares its initial Web Bot Auth protocol proposal, introducing cryptographic authentication for bot traffic using HTTP Message Signatures.
- July 1, 2025 - Cloudflare launches Pay Per Crawl in private beta, activating the HTTP 402 Payment Required status code for AI crawler monetisation and declaring the first Content Independence Day.
- August 28-29, 2025 - Cloudflare expands AI Crawl Control with customisable 402 responses, publishing data showing one operator's crawler accessing 38,000 pages for every referred visit.
- October 30, 2025 - Cloudflare publishes a registry format for bot and agent authentication, partnering with Amazon Bedrock AgentCore.
- December 20, 2025 - Cloudflare's year-end review finds AI crawlers accounted for 4.2% of HTML requests across its network in 2025, with human visitors at 43.5%.
- February 20-25, 2026 - Anthropic clarifies the separate functions of ClaudeBot, Claude-User, and Claude-SearchBot, and commits to respecting robots.txt directives.
- March 3, 2026 - Google publishes a consolidated web crawling overview page.
- March 10, 2026 - IAB Tech Lab opens the Content Monetization Protocol version 1.0 for public comment, running through April 9.
- March 20, 2026 - Google adds Google-Agent to its official crawler list, referencing an experimental web-bot-auth protocol.
- April 28, 2026 - IAB Tech Lab finalises the Content Monetization Protocol version 1, according to the guidance document.
- May 27, 2026 - IAB Tech Lab opens its Bot and Crawler Management Guidance for public comment, running through June 25.
- May 30, 2026 - Joost de Valk publishes the Website Specification, covering AI agent readiness among 128 topics.
- May 2026 - Ahrefs finds that 97% of llms.txt files received zero AI requests, despite 8.8 times adoption growth over the prior year.
- June 2026 - Automated web-page requests surpass human requests for the first time, at 57.5% of traffic, according to the guidance document's cited Cloudflare figures.
- July 1, 2026 - Cloudflare sets September 15, 2026 as the start date for default blocking of Training and Agent crawlers on ad-bearing pages for new domains, and shifts its payment model from per-crawl charging to answer-based compensation.
- July 22, 2026 - IAB Australia's Bots and Crawler Guidance and Decision Matrix is dated version 1.0, prepared by Jonas Jaanimagi.
- July 28, 2026 - IAB Australia's Discovery: AI and Search Summit takes place in Sydney, where the guidance is released.
- July 30, 2026 - IAB Australia distributes the guidance and summit recap through its weekly newsletter to members.
- September 15, 2026 - Cloudflare's default-blocking policy for Training and Agent crawlers takes effect for new domains and free-tier zones, a deadline the IAB Australia guidance directs publishers to prepare for in advance.
Related PPC Land coverage
- IAB Tech Lab opens bot management guidance for public comment - Covers the May 27, 2026 release of the IAB Tech Lab document that forms the direct technical basis for IAB Australia's matrix.
- IAB Tech Lab's CoMP spec forces LLMs to pay before they crawl - Details the March 10, 2026 opening of the Content Monetization Protocol for public comment, the licensing layer referenced throughout the guidance.
- Cloudflare exposes AI crawlers hitting sites 50000 times per visitor - Explains the July 1, 2026 Cloudflare announcement that set the September 15 default-blocking deadline embedded in the IAB Australia guidance.
- Cloudflare stops charging AI per crawl and starts paying per answer - Describes Cloudflare's July 1, 2026 shift toward answer-based compensation for publishers.
- Anthropic clarifies what its three web crawlers do - and how to block them - Background on the ClaudeBot, Claude-User, and Claude-SearchBot distinction the guidance uses as a worked example.
- Google-Agent joins the crawler list as AI browsing gets an official identity - Covers the March 20, 2026 addition of Google-Agent, one of the live-agent crawlers classified in the guidance.
- The user agent strings every SEO and site owner needs right now - A companion technical reference for the user-agent tokens catalogued across the guidance's five categories.
- Cloudflare unveils registry format for bot and agent authentication - Details the cryptographic verification approach the guidance recommends as an alternative to name-based bot identification.
- llms.txt adoption rises 8.8x but 97% of files get zero AI requests - Provides adoption data for a voluntary standard the guidance implicitly contrasts with enforceable, network-level policy.
- AI crawlers now consume 4.2% of web traffic as internet grows 19% in 2025 - Establishes the baseline traffic composition data that the guidance's 57.5% mid-2026 crossover figure builds on.
Summary
Who: IAB Australia, the trade association for online advertising in Australia, published the guidance; Jonas Jaanimagi, the organisation's Technology Lead, is credited as its author. The document is directed at publishers, brands, and technical site owners managing automated web traffic.
What: A Bots and Crawler Guidance and Decision Matrix that classifies non-human web traffic into five categories - discovery and search, AI training, live AI agents, operational infrastructure, and unknown or unverified traffic - and assigns each a recommended verdict of allow, allow with conditions, require licensing, or block.
When: The guidance is dated version 1.0, prepared July 22, 2026, and was released at IAB Australia's Discovery: AI and Search Summit on July 28, 2026, then distributed to IAB Australia's membership through its weekly newsletter on July 30, 2026.
Where: The summit took place in Sydney, Australia. The guidance itself addresses a global technical audience, drawing on data and standards from IAB Tech Lab, Cloudflare, DataDome, and multiple AI operators including Google, OpenAI, and Anthropic.
Why: The guidance responds to a structural shift the document itself quantifies: automated web traffic overtook human traffic for the first time in mid-2026, and a large share of that traffic now serves AI systems rather than conventional search indexing. With Cloudflare's own default-blocking policy for new domains set to take effect September 15, 2026, the guidance gives publishers a documented framework for deciding, crawler by crawler, what access to grant before that network-level change arrives.
Discussion