Meta-WebIndexer is the web crawler Meta Platforms uses to build and refresh the search index that supplies answers inside Meta AI. It exists because the assistant embedded in Facebook, Instagram, WhatsApp and Messenger needs live web content to respond to questions about news, prices or sports, and because Meta has spent two years trying to stop sourcing that content from Google and Microsoft. Every publisher's robots.txt file now carries, whether or not anyone edited it, a decision about whether that site can be quoted in those answers.
What Meta documents
The company's crawler page states that Meta-WebIndexer navigates the web to improve Meta AI search result quality, and that Meta analyzes online content to enhance the relevance and accuracy of the assistant. It then adds the sentence that makes the crawler commercially interesting: allowing Meta-WebIndexer in a robots.txt file helps Meta cite and link to that content in Meta AI's responses.
That is an unusually explicit trade. Most crawler documentation describes what a bot takes. This one describes what the site gets back, and by implication what it forfeits.
Two user agent strings appear in server logs. The long form reads meta-webindexer/1.1 followed by a parenthesised link to Meta's webmaster documentation. The short form is meta-webindexer/1.1 with nothing after it. The robots.txt product token is meta-webindexer, matched without regard to case under IETF RFC 9309, the 2022 standardisation of the exclusion protocol.
Meta states a preference for robots.txt over what it calls non-standard formats such as NoAI tags, and warns that changes can take up to 24 hours to register because its crawlers cache the file for that long. A publisher who blocks the crawler on a Monday morning should not read Monday afternoon's logs as evidence of non-compliance.
Verification works differently from Google's approach. Meta publishes no per-crawler JSON file of address ranges, directing operators instead to query the RADb routing registry for prefixes originating from autonomous system 32934, the number assigned to Meta. Observed ranges include 157.240.14.0/24 and 31.13.73.0/24 alongside large IPv6 allocations under 2a03:2880::/32, and Meta cautions that the addresses change often. Site owners posting in Cloudflare's community forum in March 2026 described the crawler arriving from many rotating IPv6 addresses at once, a pattern that makes address-level blocking laborious and argues for token-based rules instead.
Where it sits among Meta's crawlers
Meta documents five user agents, and the separation between them is the point. FacebookExternalHit fetches pages shared on Meta's apps to build link previews, and may bypass robots.txt when running security or integrity checks. Meta-ExternalAgent collects data for training foundation models and for indexing content directly. Meta-ExternalAds crawls for advertising and other business products. Meta-ExternalFetcher retrieves individual links at a user's request to support agentic functions, and Meta says it may bypass robots.txt for that reason. Meta-WebIndexer feeds the search index.
Because each carries its own token, a site can permit indexing while refusing training, or the reverse. Meta-WebIndexer is not listed among the crawlers that may bypass robots.txt, which places it, on paper, in the compliant category alongside Meta-ExternalAgent.
Origin and evolution
The strategic backdrop predates the token. The Information reported on October 28, 2024 that Meta was building its own AI search engine to reduce dependence on Google Search and Bing, which at the time supplied Meta AI with news, stock and sports answers. Reuters and other outlets carried the report the same day. Meta-ExternalAgent had entered the community-maintained ai.robots.txt registry three months earlier, on July 29, 2024.
Meta-WebIndexer was added to that registry on September 8, 2025. The maintainers classified its robots.txt compliance as unclear and recorded a crawl frequency of more than one request per second. The registry is a volunteer project rather than a Meta publication, and its compliance field reflects the absence of independent testing rather than evidence of misbehaviour.
Reception was shaped by what had happened a month earlier. Leaked internal documents published by Drop Site News in August 2025 showed Meta had harvested content from 6 million unique websites, its scraping systems ignoring robots.txt and reaching much of the material through content delivery networks rather than origin servers. PPC Land covered the leaked list and the circumvention it described. A crawler arriving weeks later to ask publishers to trust a robots.txt directive inherited that context.
The 2026 surge
Through most of 2026 the indexing crawler was a minor presence. Then it overtook Meta's own training crawler in June, an inversion recorded in IAB Australia's Bots and Crawler Guidance and Decision Matrix, which PPC Land examined in detail. The trade body cited the reversal as evidence that any allowlist a publisher assembles may be stale by publication. The same document noted that mixed-purpose crawling fell from roughly 49% to 33% of AI requests across the first half of 2026 as operators split composite bots into narrower tokens.
The volume followed. Promptwatch, an AI visibility platform, measured Meta-WebIndexer at roughly 2.2% of tracked AI crawler requests in mid-July 2026. The share spiked to about 23% between July 20 and 22, climbed from August 5, and reached 37.8% on August 9, according to the company's published daily figures. That is a seventeenfold increase inside a month. Promptwatch excludes Googlebot, Bingbot and link-preview fetchers from the denominator, so the figure describes composition within AI crawling rather than total site traffic.
Independent operators noticed the load before the measurement was published. A feature request filed against Anubis, an open-source bot-challenge proxy, on April 16, 2026 asked for a default rule against Meta-WebIndexer on the grounds that it was hammering the author's server. On August 6, 2026, the developer Pieter Levels posted that Meta staff had told him the company was allegedly building its own web index, and that Meta crawling had triggered load average alerts on one of his servers.
That surge landed on infrastructure already saturated. Cloudflare figures cited in the IAB Australia guidance put automated requests at 57.5% of web page requests as of June 2026, the first recorded crossover past human traffic, a threshold PPC Land reported at the time.
Why the term matters for marketers
Meta AI is not a small surface. Similarweb data covering June 2025 to May 2026 put Meta AI's growth in United States monthly active users at 435%, the fastest of the assistants measured, and noted that embedded experiences of this kind sit outside conventional chatbot traffic measurement entirely, as PPC Land documented in July 2026. Content that is not in the index cannot be cited to that audience.
The index also sits upstream of Meta's advertising machinery. Meta announced in October 2025 that interactions with its generative AI products would feed content and ad personalisation from December 16, 2025, a change PPC Land covered alongside the quarterly results. What the assistant can retrieve therefore shapes both the answers users see and the signals the ad system reads.
Concentration makes the access decision unusually consequential. HUMAN Security data reported by PPC Land put OpenAI bots at approximately 69% of observed AI-driven traffic by volume, Meta at roughly 16% and Anthropic at about 11%, leaving three companies in control of most of a typical site's automated visitors.
Limitations and disputes
Compliance is asserted rather than demonstrated. Meta does not list Meta-WebIndexer among the crawlers that may bypass robots.txt, but no independent audit confirms adherence. Third-party crawler directories disagree: some describe compliance as unverified, others state the crawler generally honours directives. That conflict is unresolved as of August 2026.
Crawl cost is real and uncompensated, and nothing in Meta's documentation offers a rate control comparable to the throttles search engines once provided. The citation promise is equally unmeasurable from the publisher's side, since Meta provides no reporting surface equivalent to Search Console.
Blocking carries its own cost. Research published in early 2026 found publishers who blocked AI crawlers lost 23.1% of monthly visits; an updated Rutgers Business School and Wharton study in April 2026 revised the effect to a 7% decline in weekly traffic within six weeks, visible in human browsing data. PPC Land reported both figures in coverage of legislative attempts to force crawler disclosure.
The separation between indexing and training is also a policy claim rather than a technical guarantee. Only Meta knows which internal systems receive the fetched bytes.
Disambiguation
Meta-ExternalAgent is the training and direct-indexing crawler. Blocking Meta-WebIndexer does not opt a site out of model training, and blocking Meta-ExternalAgent does not remove it from Meta AI search results.
Meta-ExternalFetcher retrieves single links in response to a user prompt and may ignore robots.txt. It is reactive, not systematic, and a disallow rule for the indexer will not stop it.
FacebookExternalHit generates link previews. Blocking it removes Open Graph titles, descriptions and thumbnails from links shared across Meta's apps, a distribution loss unrelated to AI.
Robots meta tags are HTML directives such as noindex placed in a page head. They share the word meta by coincidence, control indexing rather than crawling, and only take effect if the crawler is permitted to fetch the page first.
Recent developments
Cloudflare published Bot Preference Sync on August 21, 2026, a feature that rewrites robots.txt to match a site's dashboard policy and attaches disclosure conditions determining whether a crawler performing both search and training retains access to sites refusing training. PPC Land reported the launch and the four conditions attached. Meta's split-token architecture positions Meta-WebIndexer favourably under such tests, provided the separation holds in practice.
Voluntary signalling remains contested. Google's John Mueller said in July 2026 that Cloudflare's Content Signals directive changes no crawler behaviour, a position that hardened publisher scepticism toward robots.txt as a remedy of any kind.
Timeline
- July 29, 2024: Meta-ExternalAgent entered the ai.robots.txt community registry
- October 28, 2024: The Information reported Meta building an AI search engine to reduce reliance on Google and Bing
- August 2025: Leaked Meta documents showed 6 million websites harvested with robots.txt ignored
- September 8, 2025: Meta-WebIndexer added to the ai.robots.txt registry, compliance marked unclear
- December 16, 2025: Meta AI interactions began feeding ad and content personalisation
- March 21, 2026: Site owners reported meta-webindexer/1.1 crawling from rotating IPv6 addresses
- April 16, 2026: Feature request filed to add a default Meta-WebIndexer rule to the Anubis proxy
- June 2026: Meta's indexing crawler overtook its own training crawler in volume
- July 22, 2026: IAB Australia prepared crawler guidance recording the inversion
- August 6, 2026: Pieter Levels reported Meta crawling triggering server load alerts
- August 9, 2026: Meta-WebIndexer reached 37.8% of tracked AI crawler requests
- August 21, 2026: Cloudflare published Bot Preference Sync with crawler disclosure conditions
Related PPC Land coverage
- Meta leaked scraping list reveals massive content harvesting operation - The August 2025 documents showing 6 million sites harvested and robots.txt bypassed via CDN addresses.
- IAB Australia forces every crawler into one of four verdicts - The guidance recording Meta's indexing crawler overtaking its training crawler in June 2026.
- Bots overtake humans - Automated requests reaching 57.5% of web page traffic for the first time on record.
- Cloudflare blocks opaque AI crawlers from sites that disallow training - Bot Preference Sync and the disclosure conditions for dual-purpose crawlers.
- US sends 53.5% of global bot traffic, Decodo analysis finds - HUMAN Security volume shares placing Meta at roughly 16% of AI-driven traffic.
- ChatGPT loses web share to Gemini and Claude as ad penetration hits 26% - Meta AI's 435% growth in United States monthly active users and why embedded assistants escape measurement.
- Meta reports 26% revenue growth amid infrastructure spending surge - The December 16, 2025 start date for using AI chat interactions in ad personalisation.
- New York passes bill forcing AI crawlers to identify themselves to news sites - The research quantifying traffic losses for publishers that block AI crawlers.
- Google's crawler math turns against it as the open web pushes back - John Mueller's dismissal of Content Signals and the crawl-to-referral ratios behind publisher scepticism.
- The user agent strings every SEO and site owner needs right now - A reference list of AI crawler tokens and the access decisions each one governs.
- AI agents caught masquerading as humans to bypass website defenses - Spoofed user agents and distributed fetching that complicate token-based blocking.
- Cloudflare launches pay per crawl to monetize AI content access - The HTTP 402 mechanism and the roster of Meta crawlers it covers.
- Only 7.4% of Fortune 500 have an llms.txt file, study finds - How large advertisers split directives between training crawlers and search agents.
- Cloudflare CEO: Google sees 3x more web content than OpenAI through crawler monopoly - Unique-URL coverage comparisons placing Googlebot at 2.99 times Meta-ExternalAgent's reach.
- Explaining GPTBot - OpenAI's training crawler and the robots.txt token that made AI crawling addressable.
Summary
Who: Meta Platforms operates Meta-WebIndexer, with publishers, technical SEO practitioners and infrastructure providers such as Cloudflare on the receiving side. Trade bodies including IAB Australia and community registries including ai.robots.txt shape how the crawler is classified.
What: A search-indexing crawler that fetches public pages to build the index behind Meta AI's answers. It identifies itself as meta-webindexer/1.1, answers to the robots.txt token meta-webindexer, and is offered to publishers on the stated basis that allowing it makes their content eligible for citation and linking in Meta AI responses.
When: Documented by Meta and recorded in the ai.robots.txt registry on September 8, 2025, following reports in October 2024 that Meta was building an independent search index. It overtook Meta's training crawler in June 2026 and reached 37.8% of tracked AI crawler requests by August 9, 2026.
Where: Requests originate from Meta's autonomous system 32934, across IPv4 ranges including 157.240.14.0/24 and IPv6 allocations under 2a03:2880::/32. The resulting index serves Meta AI across Facebook, Instagram, WhatsApp, Messenger and the standalone application.
Why: Meta AI has depended on Google and Bing for live web answers, an arrangement that exposes query data to competitors and carries a cost. An owned index removes that dependency. For marketers, the crawler determines whether a brand's pages can be cited to a fast-growing assistant audience, while the underlying interactions also feed Meta's advertising personalisation, making a single robots.txt line consequential on both the organic and paid sides.
Discussion