Crawl budget is the number of URLs on a website that a search engine's crawler is both able and willing to fetch in a given period. Google's documentation, last updated on July 22, 2026, defines it as the set of URLs that Google can and wants to crawl. Two separate limits produce that set: one protects the server from overload, the other reflects how much the engine cares about the content. The concept exists because fetching costs money on both sides. The crawler spends computing resources, the site owner pays for bandwidth and processing, and neither can be unlimited.
How the allowance is set
Google splits the calculation into two parts. The crawl capacity limit caps the number of simultaneous connections and the pause between fetches, according to Google's documentation, rising or falling with response times, server errors and Google's own resources. Crawl demand varies by crawler; for Googlebot it depends on perceived inventory, popularity and staleness. Capacity is a ceiling and demand decides how much of it gets used. Google said in its 2017 post, as reported by Search Engine Roundtable, that Googlebot will not crawl without sufficient demand even when capacity is free.
The pipeline runs in stages. Discovery comes from links, sitemaps and Search Console submissions, and recrawl timing follows observed change patterns. Google's overview of March 3, 2026 said the homepage of a breaking-news site may be recrawled every few minutes, while pages unchanged for years may wait a month. Retail product pages are revisited often to reflect prices, promotions and stock.
Fetching HTML is only the first step. The page goes to Google's Web Rendering Service (WRS), a headless browser that runs JavaScript, and every referenced resource is fetched separately. A December 2024 Google document stated that resource management can significantly affect a site's budget, and that the WRS caches JavaScript and Cascading Style Sheets (CSS) files for 30 days regardless of HTTP caching headers. Cache-busting parameters on those files therefore carry a cost.
Size limits apply per fetch. On February 6, 2026 Googlebot's cap for a single URL fell from 15MB to 2MB, with Portable Document Format (PDF) files kept at 64MB. A later clarification said Googlebot stops at the exact cutoff and treats the partial file as complete, and that each rendering resource has its own byte counter. Analysis by Dave Smart suggested fewer than 0.01% of sites are affected.
Who manages it
The work sits on the sell side: publishers, retailers and marketplaces, through search engine optimisation (SEO) and engineering teams, with hosting providers and content delivery networks (CDNs) in the background. The levers are mostly indirect. Servers signal strain through slow responses and 5xx errors. Robots.txt disallow rules keep crawlers away from low-value paths, though Google's specification says crawl-delay is not supported. Sitemaps with accurate last-modified dates, 404 or 410 codes for removed pages and HTTP 304 responses all cut wasted fetches, according to Google's documentation. Server logs and the Crawl Stats report in Search Console show what was requested.
For emergencies, Google's guidance on reducing crawl rate is to return 500, 503 or 429 status codes, with a Retry-After header on the last two, for no more than one to two days. Beyond that, URLs may drop from the index. Google also states that an increase in crawl rate cannot be requested.
Buy-side teams hold no controls here. Yet the landing pages and product pages that receive paid traffic sit on the same servers, and how quickly new ones reach organic results depends on this system.
Bing offers a different tool. Its Crawl Control, described in a December 2018 post by Nikunj Daga, lets webmasters set an hourly crawl rate, slower in peak business hours and faster off-peak.
Origin and evolution
Search Console already offered a crawl rate setting, which by 2023 had existed for more than a decade, before the vocabulary was formalised. On January 16, 2017, Google's Webmaster blog introduced the terms crawl rate limit, crawl demand and crawl budget, with Gary Illyes as author, according to Search Engine Journal. It named faceted navigation, soft errors, hacked pages, infinite spaces and low-quality content as drains.
Measurement came next. In November 2020 Google released a rebuilt Crawl Stats report showing total requests, download size and response time, split by host status, response code, file type, purpose and Googlebot type.
Control then moved the other way. Google announced in November 2023 that the crawl rate limiter tool, available for over a decade, would go. It was removed on January 8, 2024. Illyes said that its usefulness had dissipated, that it took more than a day to take effect and that it was rarely used. Googlebot, Google said, already responds to server signals.
Late 2024 brought a documentation push. A November 1, 2024 change required sites with separate mobile HTML to keep link parity with desktop, aimed at sites above one million unique pages or 10,000 pages updated daily. On December 18, 2024 Google published faceted navigation guidance warning that filter parameters can create nearly infinite URL variations.
Why it matters to marketers
Large catalogues generate URLs faster than any crawler wants them. Colour, size and sort parameters multiply product listings, and each combination consumes fetches. Google's documentation sets the threshold for concern at sites with more than 1 million pages changing weekly, or more than 10,000 pages changing daily, plus sites with many URLs marked "Discovered - currently not indexed".
Timing is the practical stake. At Search Central Live Deep Dive Europe in Barcelona on October 2, 2026, Illyes presented internal figures: a new URL takes about 20 hours to be found, a known URL about 30 days to refresh, with slowest cases running to weeks or never. Sitemap processing typically took about 24 hours. For a seasonal campaign page, those intervals matter.
Quality feeds into the same system. John Mueller said in an episode released on October 1, 2026 that crawl demand is often based on perceived quality, so a technically valid sitemap offers no protection for a site Google considers weak. Martin Splitt explained in August 2024 that the "Discovered - currently not indexed" label means Google has found a URL but not yet fetched it, often because of queue position, server strain or low perceived value.
Limits and disputes
Most sites fall below Google's thresholds, so advice to optimise the allowance is often misdirected. Disputes concentrate on opacity. Google publishes no per-site number, and the Crawl Stats report records requests already made rather than an entitlement.
Illyes said in a 2024 podcast that higher crawl frequency does not by itself signal higher quality, and that hosting companies sometimes blame Google for faults inside their own infrastructure. Mueller added that responses of several seconds sharply cut the number of pages fetched.
Independent observers cannot easily audit Google's side. From August 8, 2025, sites on Vercel, WP Engine and Fastly saw crawl rates fall, with Vercel's chief technology officer Malte Ubl documenting a 30% global drop. Mueller confirmed on August 28 that the cause was on Google's side and had been resolved. Rankings held steady throughout, which indicated that Google relied on cached information.
Controls remain blunt. A consultant, Javier Lorente Murillo, reported that over half of Googlebot's requests to his classifieds site returned 404. Illyes replied that the expiry tag only works once the page is crawled again. Internal search pages are another sink: Mueller described infinite crawl space when millions of parameter combinations are indexable, a reason Google still recommends blocking them even after dropping a 2007 guideline on July 31, 2026.
Not the same as
Crawl rate. The speed component, meaning connections and pauses. It is one half of the capacity limit, not the whole allowance.
Indexing. Retrieval and storage are separate steps. A page can be crawled and never stored, or discovered and never fetched.
Crawl-delay. A robots.txt line asking for pauses between requests. Google ignores it; Bing's hourly Crawl Control serves a similar purpose.
Sitemap hints. Priority and change-frequency fields were deprecated; last-modified dates remain useful. A single sitemap file is capped at 50,000 URLs or 50MB uncompressed, and nothing in the format grants extra crawling.
Recent developments
Artificial intelligence (AI) crawlers have widened the cost argument. Cloudflare data for July 2025 showed training taking 79% of AI bot activity and, per Cloudflare, crawl-to-referral ratios of 38,065 to 1 for Anthropic, 1,091 to 1 for OpenAI and 5.4 to 1 for Google. Those ratios compare pages fetched with visits sent back.
Individual sites report the load. Kinsta recorded 3.75 million requests to one WooCommerce cart page in 24 hours, with ClaudeBot the largest source among the bots measured. "Every request is real work," Kinsta chief technology officer Daniel Pataki said. PatronView's founder Nick Gray blocked two AI crawlers after his server logged 2.5 million requests in a week against 5,977 analytics pageviews. A Cloudflare and ETH Zurich study said such traffic strains CDN cachingbuilt around human patterns, because unique AI requests push popular pages out of the cache. In an October 2026 report, the Wikimedia Foundation said bots generated 65% of its most resource-consuming traffic in 2025, and that bandwidth use had risen 50% since 2024.
Google's allowance adjusts to server health. The incidents above describe sites absorbing load with no comparable throttle.
Timeline
- January 16, 2017: Google's Webmaster blog defines crawl rate limit, crawl demand and crawl budget
- December 2018: Bing describes Crawl Control for setting hourly crawl rates
- November 2020: Google releases the rebuilt Crawl Stats report
- November 2023: Google announces the crawl rate limiter tool will be retired
- January 8, 2024: The crawl rate limiter tool is removed
- August 2024: Martin Splitt explains the "Discovered - currently not indexed" status
- November 1, 2024: Google requires link parity on separate mobile pages for large sites
- December 3, 2024: Google publishes a document on rendering, resources and the 30-day cache
- December 18, 2024: Google publishes faceted navigation guidance
- August 8, 2025: Crawl rates fall for sites on several hosting platforms
- August 28, 2025: John Mueller confirms a Google-side issue, now resolved
- February 6, 2026: Googlebot's per-URL limit falls from 15MB to 2MB
- March 3, 2026: Google publishes its overview of how crawling works
- July 22, 2026: Google's crawl budget documentation is last updated
- October 2, 2026: Gary Illyes presents discovery and refresh timing figures in Barcelona
Related PPC Land coverage
- Explaining Googlebot - Architecture, fetch limits and the crawl process behind the allowance.
- Google slashes web crawl limit by 86.7% - The February 2026 drop from 15MB to 2MB per URL.
- Google details comprehensive web crawling process - The December 2024 document on rendering, resources and the 30-day cache.
- Google's secret crawl logic, finally explained in one page - Google's March 2026 overview of crawl frequency and scheduling.
- Google rewrites Googlebot's rulebook - How the 2MB cutoff operates and how Googlebot is described.
- Google Search Console crawl rate limiter tool to be deprecated - The November 2023 announcement ending the manual limiter.
- Managing faceted navigation URLs - Google's December 2024 guidance on filter parameters.
- Google updates mobile site requirements - Link parity for large sites with separate mobile pages.
- Google crawl rate declines affect multiple hosting platforms - The August 2025 fall in crawling and Google's explanation.
- Google says new pages take about 20 hours to be found - Illyes's October 2026 discovery and refresh timing figures.
- Google may skip sitemaps on sites it deems low quality - Mueller on crawl demand, quality and sitemap fetching.
- Google explains 'Discovered - Currently Not Indexed' - Splitt's account of why URLs wait in the queue.
- Google search team explores web crawling challenges - A 2024 podcast on crawl frequency, server speed and parameters.
- Google's expiry tag forces a re-crawl - Why the tag does not reduce 404 crawl waste on fast-turnover sites.
- Google drops 2007 rule requiring blocked internal search pages - Mueller on infinite crawl space from search result URLs.
- Explaining indexing - The distinction between retrieval and storage.
- AI bots hammered WordPress cart pages 3.75M times in a day - Kinsta's data on crawler load against uncacheable pages.
- PatronView blocks Amazon's AI crawler after 117,000 daily page reads - A small publisher's account of crawler volume against analytics.
- Cloudflare and ETH Zurich say AI bots are breaking the web's cache layer - Research on cache eviction caused by AI crawling.
- Wikimedia says OpenAI agents may have contributed to partial outage - Bot shares of resource-heavy traffic at Wikimedia.
Summary
Who: Search engines such as Google and Bing set the allowance for their own crawlers; publishers, retailers and their engineering and SEO teams influence it. Gary Illyes of Google defined the term in a 2017 post.
What: The set of URLs a crawler can fetch without straining a server and wants to fetch given popularity, freshness and perceived quality. It matters chiefly to sites with at least 10,000 daily-changing or 1 million weekly-changing pages.
When: Defined on January 16, 2017; the manual limiter was removed on January 8, 2024; the per-URL fetch cap fell to 2MB on February 6, 2026; Google's documentation was last updated on July 22, 2026.
Where: Inside each search engine's crawling systems, observed through server logs and the Crawl Stats report in Search Console, and shaped by robots.txt, sitemaps, status codes and server speed.
Why: Crawling costs both sides money, and a site's useful pages need to be found and refreshed before low-value URLs absorb the fetches. Google says most sites need not worry, while AI crawler traffic has renewed the argument over who bears the cost.
Discussion