Cloudflare declared AI Search generally available on October 1, 2026, and set November 1, 2026 as the date on which charges begin for the managed retrieval service. The release adds native image embeddings and raises the file size cap, according to a post on the company's blog by Gabriel Massadas, Nelson Duarte and Ashish Vinodkumar.
In Short
Cloudflare has finished testing a tool that lets websites and apps search their own documents and images, and it will start charging for it on November 1, 2026. The change affects developers and companies that run site search or internal knowledge search on Cloudflare, because the price depends on how much content is loaded, how much is stored and how many searches are run. Small projects stay free under monthly limits, while heavier use is metered, and scanned PDFs and pictures can now be searched as well.
What the product does
Cloudflare describes AI Search as a product that "combines Workers AI, Vectorize, R2, and Browser Run into a fully managed index and retrieval pipeline". According to Cloudflare, the service launched more than a year ago, and developers have used it for everything from searching internal documentation to powering search on public websites. The company also runs its own blog and developer docs search on it.
The product has carried another name. Cloudflare introduced AutoRAG in open beta in April 2025 as a fully managed retrieval-augmented generation pipeline, and AI Index, a private beta from September 2025, processes new and updated pages in real time using technology from AI Search, previously known as AutoRAG.
How a query moves through the system
Cloudflare lays out the query path in sequence. The query is optionally rewritten and then embedded. Image queries are embedded directly by multimodal models, or captioned first when the embedding model handles text only. Vector search and keyword search run in parallel, and the results are fused and optionally reranked. The top chunks are either returned as they are or passed to a generation model that writes an answer.
Multimodal retrieval
The earlier handling of images was, in Cloudflare's own word, naive. The system ran object detection, generated a caption and embedded that text, which left images searchable only through whatever the caption happened to capture. The new version does both: it embeds image pixels directly for visual retrieval and keeps captions for textual understanding.
To hold down the cost of the larger representations, AI Search uses Matryoshka Representation Learning, a method that lets smaller embeddings retain useful information so that storage stays manageable and search stays fast, according to Cloudflare. Native multimodal retrieval is available with the Qwen3-VL-Embedding model. At query time the system checks whether the instance's embedding model supports images. If it does, a query image is embedded by that model and lands in the same vector space as the indexed images and text.
A text-only embedding model still accepts an image query. In that case AI Search converts the image to text with ToMarkdown and searches on the resulting caption. Cloudflare says this gives every model basic multimodal support, while models with native image support receive the full visual signal.
The post illustrates the difference with a photograph of a small bird. The caption reads "A small bird with black-and-white markings perched among golden fruit and green leaves". Native retrieval, according to Cloudflare, can match details the caption omits, among them the geometry of the white eyebrow stripe, the yellow-green plumage, the mix of smooth and weathered fruit and the shallow-focus composition. Cloudflare puts the limitation of captions this way: "The caption is a compressed interpretation of the image." The company lists product discovery, screenshot matching, charts, diagrams and scanned documents as uses where color, texture, composition or spatial relationships matter.
Larger files and OCR
AI Search now accepts text files (Markdown, HTML, CSV, JSON and similar) and PDFs of up to 10 MiB, up from 4 MiB. Many PDFs are in fact scanned images with no extractable text. For those, a switch turns on OCR, and the service reads the text from each page before chunking and embedding it. OCR is open to every account and is billed under the new pricing as image processing ingestion tokens.
Pricing
Cloudflare announced preview pricing during its August 2026 Agents Week. With general availability, billing goes live on November 1, 2026, and the company says it will send a reminder email before that happens.
The structure has three charged elements: the content ingested, the data stored and the queries run. Parsing, chunking, embedding with Workers AI models, keyword indexing and reranking sit inside those charges. Cloudflare says there are no instance hours, no capacity units and no monthly minimums to size in advance. Ingestion is priced at one rate per token whichever Workers AI embedding model is selected, and tokens are counted the same way for every model. Moving from a text-only embedding model to a multimodal one therefore leaves the ingestion rate unchanged, unless images are also processed, which adds a fee.
| Item | Rate | Free monthly allotment, all Workers plans |
|---|---|---|
| Base ingestion | $0.75 per 1M tokens | 5M tokens |
| Image processing (add-on) | +$0.50 per 1M tokens | 5M tokens |
| Stored data | $2.00 per GB-month | 10 GB |
| Semantic queries (hybrid and vector search) | $0.75 per 1k queries | 1,000 queries |
| Full-text queries | $0.10 per 1k queries | 1,000 queries |
| Embedding and reranking | Free with select Workers AI models; third-party models billed separately | Not applicable |
The two 5M token allotments are a single monthly pool covering any supported file type, text and images alike, according to a footnote in the post.
One change separates the final schedule from the preview. The free allotment now holds 1,000 semantic queries and 1,000 full-text queries, replacing a shared pool of 2,000. Arithmetic on the published rates puts one million semantic queries at $750 and one million full-text queries at $100 before any free allotment is applied. The post does not say which Workers AI models count as "select" for free embedding and reranking.
Why the release matters to marketing and publishing teams
Retrieval layers can sit behind on-site search, product discovery and question-answering features on publisher and retail sites, and Cloudflare names website search among the uses of AI Search. A per-token, per-gigabyte and per-query schedule makes the cost of such features computable from content volume and traffic, which differs from capacity-based plans where a fixed unit has to be chosen first. Cloudflare presents that predictability as a design goal, stating that the pricing allows a bill to be estimated before a single file is indexed.
Spend controls have appeared elsewhere in Cloudflare's AI lineup. In June 2026 the company added dollar-denominated spend limits to AI Gateway in public beta, with caps by team, user or model.
The announcement also carries a limit of its own. The post gives no figures on retrieval accuracy, latency or customer deployments, so the practical effect of native image embeddings against the caption-only approach rests on Cloudflare's description and an example of one photograph. Independent measurements are not part of the material.
What Cloudflare says comes next
Three items are listed, none with a date. First, an ingestion pipeline for full video and audio processing, so that customers can search rich media assets. Second, a refactor of the keyword search engine so that it scales better, particularly for large data stores where the current implementation has limits. Third, better and simpler ways to enable AI Search and create indexes for websites already running on Cloudflare, which Cloudflare says would help AI agents discover, explore and consume content more easily and efficiently. Cloudflare describes these as follow-up announcements still to come.
Timeline
- April 7, 2025: Cloudflare introduces AutoRAG, a managed retrieval-augmented generation pipeline, in open beta.
- September 26, 2025: Cloudflare launches AI Index in private beta, built on technology from AI Search, previously known as AutoRAG.
- June 5, 2026: Cloudflare releases dollar-denominated spend limits for AI Gateway in public beta.
- August 2026: Cloudflare announces preview pricing for AI Search during Agents Week.
- October 1, 2026: Cloudflare declares AI Search generally available, adding native image embeddings, OCR for scanned PDFs and a 10 MiB file limit.
- November 1, 2026: Billing for AI Search goes live, with a reminder email promised beforehand.
Related PPC Land coverage
- Cloudflare unveils major AI Agent development tools: April 2025 coverage of the AutoRAG open beta, alongside the Outerbase acquisition and the general availability of Workflows.
- Cloudflare launches AI Index for website content discovery: September 2025 report on the private beta that lets website owners expose structured access to content for AI platforms, built on AI Search technology.
- Cloudflare AI Gateway now caps runaway AI bills with dollar budgets: June 2026 coverage of spend limits by team, user or model and a closed beta for identity-driven budgets.
Summary
- Who: Cloudflare, with the post written by Gabriel Massadas, Nelson Duarte and Ashish Vinodkumar. The audience is developers on Workers plans and organizations that run site or document search on Cloudflare.
- What: AI Search, a managed index and retrieval pipeline built from Workers AI, Vectorize, R2 and Browser Run, became generally available with native image embeddings through Qwen3-VL-Embedding, OCR for scanned PDFs, a 10 MiB file limit (up from 4 MiB) and a usage-based price list.
- When: The announcement came on October 1, 2026. Billing begins on November 1, 2026, and preview pricing was first announced during Agents Week in August 2026.
- Where: The service runs on the Cloudflare developer platform and is open to every account, with a free monthly allotment on all Workers plans.
- Why: According to Cloudflare, images carry details that captions lose, scanned PDFs hold text that plain extraction misses, and a pricing model based on ingestion, storage and queries lets a bill be estimated before any file is indexed. Cloudflare also plans video and audio ingestion, a keyword search refactor and simpler indexing for sites on its network.
Discussion