Anthropic today released Claude Haiku 5.5, a small model priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts of up to 100,000 tokens, against $1.00 and $5.00 for Haiku 4.5, according to the company's release page. The company also halved the price of cache reads on Claude Sonnet 5.5 and added monthly API credits for Max and Team subscribers.
In Short
Anthropic, the company that makes Claude, today put out a smaller and cheaper AI model called Haiku 5.5, and it made one of its mid-sized models cheaper to run as well. That matters to any business whose software sends Claude thousands of short, repetitive jobs a day, because each job now costs a fraction of what it did on the earlier small model. Anthropic says the new model suits quick work such as sorting and summarizing, while its bigger models stay the ones meant for harder tasks.
A small model with a lower rate card
Anthropic describes Haiku 5.5 as the "cheapest, fastest, and most capable small model" it has released, a characterisation that rests on the company's own evaluations. According to Anthropic, the model is designed for "high-volume, cost-sensitive tasks": summaries, compactions, database queries and classification requests, along with live customer support and browser use, where response time carries more weight. The company also positions it as a subagent, a model handed one slice of a larger job, working alongside Opus 5.5 and Sonnet 5.5 on coding work.
The speed claim carries a footnote. According to Anthropic, Haiku 5.5 is its fastest model at each model's standard speed, although it runs less quickly than the Opus models in Fast Mode. The page gives no latency figures of its own beyond one customer's result, covered below.
Availability is broad from the first day. Haiku 5.5 is available now on all platforms, Anthropic says, including Amazon Web Services, Google Cloud and Microsoft Azure. Developers on the Claude Platform call it with the identifier claude-haiku-5-5, and a migration guide covers the move from earlier models.
The release closes a three-model cycle. Anthropic released Claude Opus 5.5 on September 22, 2026, at $4 per million input tokens and $20 per million output tokens, according to AI Business, and said at the time that Sonnet 5.5 and Haiku 5.5 would follow within weeks, according to TechJack Solutions. Sonnet 5.5 arrived on September 28, according to the LLM Gateway model listing. Haiku 5.5 therefore comes nine days after Sonnet and fifteen days after Opus.
Benchmark results
Anthropic published results across seven benchmarks, with Humanity's Last Exam shown both with and without tools. The page labels the Sonnet 5.5 column as a reference point rather than a competitor. Details of how the evaluations were run sit in the Haiku 5.5 System Card, which was not part of the material reviewed for this article, and all figures are vendor-reported.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (knowledge work) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (knowledge work) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1, offline subset (computer use) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam, no tools | 45.9% | 10.2% | not reported | 56.9% |
| Humanity's Last Exam, with tools | 57.4% | 18.7% | not reported | 64.5% |
| Terminal-Bench 4.0 (agentic coding) | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 Main (agentic coding) | 46.4% | not reported | 42.4% | 52.1% (Xhigh) |
| Chartography, no tools (visual reasoning) | 46.4% | 6.4% | 29.1% | 61.6% |
Against its predecessor, the gains are large on the agent-style tests. OSWorld 2.1, which according to Anthropic measures how well agents can operate a real computer to finish long, multi-step tasks, moves from 15.7% to 72.4%, a little over four and a half times. Chartography rises from 6.4% to 46.4%, more than sevenfold, and Terminal-Bench 4.0 from 0.0% to 39.2%. Terminal-Bench 4.0 measures, in Anthropic's description, complex, multi-step professional tasks within a command-line interface.
GPT-6 Luna, an OpenAI model according to AI Business, is the only third-party entry. Haiku 5.5 leads it on all six benchmarks where Luna has a result. The margin runs from 4.0 points on FrontierCode (46.4% against 42.4%) to 23.5 points on OSWorld (72.4% against 48.9%). Luna has no entry for Humanity's Last Exam.
Sonnet 5.5 stays ahead of Haiku 5.5 on every row. The widest gap is on Terminal-Bench 4.0, where Sonnet 5.5 scores 70.6% against 39.2%, a difference of 31.4 points; the narrowest is FrontierCode at 5.7 points, with the Sonnet figure recorded at its Xhigh setting. Anthropic draws the line itself: Sonnet 5.5 and Opus 5.5 "remain better choices" for complex agentic coding, while Haiku 5.5 suits "more narrowly scoped tasks" that earlier Claude versions would have made cost-prohibitive, such as compaction, summarization or subagent work.
Effort settings
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, Anthropic says, so users can choose whether to optimize for cost or for intelligence, as with the company's other models. The page offers charts for OSWorld 2.1, GDPval-AA and Humanity's Last Exam, each plotting score against cost per attempt, with settings labelled Low, Med, High, Xhigh and Max. The OSWorld view is the one displayed in the copy reviewed. On its logarithmic cost axis, running from $0.05 to $5, the Haiku 5.5 curve ends just before the Sonnet 5.5 curve begins and sits above the GPT-6 Luna curve where the two overlap in cost.
The table does not state which effort setting produced each Haiku 5.5 score, except for the labelled Xhigh entry on Sonnet 5.5.
Pricing in detail
Anthropic splits Haiku 5.5 pricing in two by prompt length. Prompts of up to 100,000 tokens get the lower rates, and the company says these account for around 90% of requests to the previous Haiku model.
| Price per 1 million tokens | Haiku 5.5, prompts up to 100k | Haiku 5.5, prompts over 100k | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Cache reads | $0.01 | $0.05 | $0.10 | $0.10 |
| Cache writes | $0.125 | $0.625 | $1.25 | $2.50 |
| Input tokens | $0.10 | $0.50 | $1.00 | $2.00 |
| Output tokens | $0.50 | $2.50 | $5.00 | $10.00 |
Anthropic's footnote states that Haiku 5.5 is priced 90% lower than Haiku 4.5 for requests up to 100,000 tokens and 50% lower above that threshold. The headline figure, that the model costs around 75% less to run on average, also accounts for a change in how many tokens a given piece of work consumes. Haiku 5.5 has an updated tokenizer, similar to those of Sonnet 5.5 and Opus 5.5, and uses "slightly more tokens per task", according to Anthropic. The company does not say how many more.
The table also exposes the structure of the price list. Output tokens cost five times the input rate in every column. Cache writes cost 1.25 times input in every column. Cache reads cost one-tenth of input on both Haiku models, at $0.01 against $0.10 on Haiku 5.5 for shorter prompts. Sonnet 5.5, after today's change, now charges one-twentieth of its input rate for cache reads, a deeper relative discount than either Haiku model offers.
Illustrative arithmetic from the list prices, before any tokenizer effect, shows the spread: a request with 50,000 input tokens and 1,000 output tokens costs $0.0055 on Haiku 5.5, $0.055 on Haiku 4.5 and $0.11 on Sonnet 5.5. No cache hits are assumed, and any hits would lower each figure. Across one million such requests, that is $5,500, $55,000 and $110,000. Haiku 5.5's input and output rates for shorter prompts are one-twentieth of Sonnet 5.5's.
Sonnet 5.5 cache reads and subscriber credits
The Haiku release came with two further changes, both taking effect around today.
First, Anthropic lowered the price of cache reads on Sonnet 5.5 by 50%, from $0.20 to $0.10 per million tokens. Because cache reads make up a large share of token consumption in agentic work, the company estimates the cut reduces the cost of Sonnet 5.5 on most agentic tasks by around 20%. A Terminal-Bench 4.0 chart on the page plots accuracy against cost per attempt for Haiku 5.5, Haiku 4.5 and Sonnet 5.5 at both the $0.10 and $0.20 cache-read prices, on a logarithmic axis from $0.50 to $10.
Second, a new monthly API credit rolls out this week to all Max and Team subscribers for use on the Claude Platform, Anthropic says. Max 5x subscribers receive $100 a month and Max 20x subscribers $200. Team subscribers receive up to $500, pooled across users. The credits can be spent on any Anthropic model and are intended for experimenting with tools, apps and agents that call the API. Anthropic points to a Help Center article for terms, which was not reviewed.
For developers, the company is updating its Claude Python and TypeScript SDKs to add computer use and browser use in beta, an area where Anthropic describes Haiku 5.5 as well suited given its speed, capability and price.
Safety and safeguards
According to Anthropic, Haiku 5.5 shows major improvements across almost all of its alignment evaluations relative to Haiku 4.5, with far fewer instances of misaligned behavior and a lower willingness to cooperate with misuse. The model's system card holds the detail.
Cybersecurity safeguards are more restrictive than those on Haiku 4.5 but somewhat less restrictive than those applied to other recent Anthropic models, the company says. They permit a wider range of defensive tasks than the safeguards on Sonnet 5.5 while still blocking penetration testing and other techniques more likely to be used by attackers. Biology safeguards match those on Sonnet 5, Sonnet 5.5 and Opus 5: research biology questions are allowed, and requests Anthropic judges likely to cause harm are restricted. Organizations doing wider-ranging biology and cyber work can apply to the Life Sciences Verification Program and the Cyber Verification Program.
Customer feedback
The page lists comments from six customers: Asana, HubSpot, AlphaSense, Box, Rogo and Cognition. Asana's was displayed in full in the copy reviewed. According to Aaron Vinh, Staff Software Engineer at Asana, the company ran Haiku 5.5 through its evaluation suite for AI Teammates, its AI agent product, covering tasks such as triaging bugs, setting up projects and searching large portfolios for high-risk or overdue work. Compared with the model Asana uses today, he reported "over a 30% reduction in latency for task completions" and up to 2.5 times faster inference per agent turn, summing it up as "a noticeably snappier experience". The comparison model is not named, and the figures are the customer's own test results as published by Anthropic.
Why the figures matter beyond developers
For marketing and ad tech, the relevance lies in cost per unit of automated work. Small-model pricing sets the floor for any product that makes many model calls: tagging, report summaries, support chat, browser-driven tasks. Anthropic's listed use cases overlap with those chores, and customers on its page include a CRM platform (HubSpot) and a work-management vendor (Asana).
Tokens, the billing unit, are only one line in an agent's bill. PPC Land reported on October 4 that Cloudflare's web search for AI agents costs $0.25 to $7 per 1,000 requests, depending on which of three providers handles the query, on top of the model charges. A cheaper model lowers one component; retrieval, tooling and orchestration costs sit elsewhere.
Pricing pressure at the small end is not confined to Anthropic. Google is cutting two of three Gemini models for free users from October 9, leaving free accounts with Flash-Lite only and moving Deep Think to the AI Pro plan at $19.99 a month. Anthropic's own comparison table sets Haiku 5.5 against an OpenAI small model rather than only against its own line-up, which places the release inside a contest over small-model economics.
The subagent framing connects to how agentic AI products are being assembled: a larger model plans, smaller ones execute delegated pieces, and the cost of the executing tier dominates at volume. Anthropic's SDK beta for computer use and browser use points to agents acting inside web interfaces, a category that MLCommons catalogued for privacy risk, listing 25 vectors across five domains in the first version of its Agent Privacy Risk Taxonomy. That taxonomy concerns agents in general and does not mention Haiku 5.5.
Open questions remain. The 75% average depends on the token-count change Anthropic did not quantify, and every benchmark figure comes from the vendor. Independent evaluation of Haiku 5.5 against the small models it is compared with has not yet appeared in the material reviewed.
Timeline
- September 22, 2026: Anthropic releases Claude Opus 5.5, priced at $4 per million input tokens and $20 per million output tokens, and says Sonnet 5.5 and Haiku 5.5 will follow in coming weeks.
- September 28, 2026: Claude Sonnet 5.5 is released at $2 per million input tokens and $10 per million output tokens.
- October 1, 2026: MLCommons publishes version 0.1 of its Agent Privacy Risk Taxonomy, with 25 risk vectors across five domains.
- October 4, 2026: PPC Land reports Cloudflare's web search pricing for AI agents, $0.25 to $7 per 1,000 requests, and Google's plan to cut two of three Gemini models for free users.
- October 7, 2026 (today): Anthropic makes Claude Haiku 5.5 available on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure, and halves Sonnet 5.5 cache-read pricing to $0.10 per million tokens.
- Week of October 7, 2026: Monthly API credits of $100 (Max 5x), $200 (Max 20x) and up to $500 pooled (Team) begin rolling out.
- October 9, 2026: Google's Gemini changes for free accounts take effect.
Related PPC Land coverage
- Cloudflare web search for AI agents costs $0.25 to $7 per 1,000 requests - Per-request pricing for agent web search in AI Gateway, routed to three providers with no markup (October 4, 2026).
- Google cuts two of three Gemini models for free users from October 9 - Changes to free and paid Gemini plan tiers, with limits of 2x for AI Plus and 4x for AI Pro (October 4, 2026).
- MLCommons catalogues 25 privacy risks posed by AI agents - Version 0.1 of an agent privacy risk taxonomy, with benchmarks targeted for 2027 (October 4, 2026).
Summary
Who: Anthropic, the developer of Claude, with Max and Team subscribers, API developers and enterprise customers such as Asana, HubSpot, AlphaSense, Box, Rogo and Cognition named on the release page.
What: Claude Haiku 5.5, a small model priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, with an adjustable effort setting; a 50% cut in Sonnet 5.5 cache-read pricing to $0.10 per million tokens; new monthly API credits for Max and Team plans; and SDK beta support for computer use and browser use.
When: Today, October 7, 2026, with the subscriber credits rolling out this week.
Where: The Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure, under the model identifier claude-haiku-5-5.
Why: According to Anthropic, to give high-volume, cost-sensitive and speed-sensitive workloads a cheaper and faster model, to lower the cost of Sonnet 5.5 on agentic work by around 20%, and to support developers building agents on the Claude Platform.
Discussion