On October 7, 2026, Anthropic released Claude Haiku 5.5, which it calls its cheapest, fastest and most capable small model. The headline is price: Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, a tenth of Haiku 4.5's rates. Anthropic also added adjustable effort to the Haiku line for the first time and cut Sonnet 5.5's cache-read price in half. This article lays out the confirmed numbers, shows what they mean per task, and explains where a small model fits next to Sonnet 5.5, Opus 5.5 and rival small models.
Anthropic
GeminiWhat Claude Haiku 5.5 is
Per Anthropic's launch page and its model overview, the model ID is claude-haiku-5-5. It has a 1 million token context window and up to 128,000 output tokens, accepts text and images, and returns text. Adaptive thinking is on by default, steered by an effort parameter whose default for this model is medium. The launch page shows results at Low, Med, High, Xhigh and Max effort. Anthropic describes it as the first Haiku-class model with adjustable effort, and positions it for classification, routing, extraction and subagent work.
It is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic's lifecycle table lists retirement as not sooner than October 7, 2027.
One detail changes the cost math: Haiku 5.5 uses the newer tokenizer shared by Claude 4.7 and later, which counts about 30% more tokens for the same text than Haiku 4.5 does. Anthropic's footnote says the model is 90% cheaper than Haiku 4.5 for requests up to 100,000 tokens, 50% cheaper above that, and about 75% cheaper on average. It also notes Haiku 5.5 uses slightly more tokens per task.
The pricing, in one table
| Model | Input | Output | Cache read | Cache write (5 min) |
|---|---|---|---|---|
| Haiku 5.5, prompt up to 100K | $0.10 | $0.50 | $0.01 | $0.125 |
| Haiku 5.5, prompt over 100K | $0.50 | $2.50 | $0.05 | $0.625 |
| Haiku 4.5 | $1.00 | $5.00 | $0.10 | $1.25 |
| Sonnet 5.5 | $2.00 | $10.00 | $0.10 | $2.50 |
| Opus 5.5 | $4.00 | $20.00 | $0.20 | $5.00 |
USD per million tokens, from Anthropic's pricing page. The Batch API takes 50% off input and output.
Two structural points stand out. First, Haiku 5.5 is the only current Claude model with length-tiered pricing: prompts over 100,000 tokens pay five times the base rate, while the other 4.6-and-later models, per the pricing page, include their full 1M window at one rate. A small model is therefore cheapest for short, high-volume calls, and the advantage narrows for long-document work.
Second, Sonnet 5.5 and Opus 5.5 now price cache reads at 5% of input rather than the usual 10%. Anthropic says this halves Sonnet 5.5's cache-read price from $0.20 to $0.10 and lowers its cost on most agentic tasks by about 20%. Our Sonnet 5.5 launch coverage listed the earlier $0.20 rate.
Cost per task: what the price gap looks like
List prices are hard to compare in the abstract, so the chart applies them to one illustrative request: 10,000 input tokens and 1,000 output tokens, uncached. This is our own arithmetic, not a vendor benchmark.
List-price cost of one 10,000-in, 1,000-out request
Our arithmetic on published standard list prices, uncached and not batched. Tokenizers differ by vendor, so the same text will not count identically across models.
That request costs $0.0015 on Haiku 5.5, $0.015 on Haiku 4.5, $0.03 on Sonnet 5.5 and $0.06 on Opus 5.5. Adjusting Haiku 5.5 for the roughly 30% token inflation (13,000 in, 1,300 out) gives about $0.00195, which is still around 87% below Haiku 4.5 on the same text. That is in line with Anthropic's 75% average once longer prompts are included.
Caching widens the gap for agents. Re-reading a 100,000-token cached prompt costs about $0.001 on Haiku 5.5 and about $0.01 on Sonnet 5.5 (cache-read rates times 100,000 tokens). An agent that resends a large context at every step feels that difference on each loop.
Benchmarks: the gap to Sonnet 5.5
The numbers below are Anthropic's own, published on the launch page, with GPT-6 Luna and Sonnet 5.5 shown for reference. They are vendor-reported and we have not independently reproduced them.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| OSWorld 2.1, offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| Humanity's Last Exam, no tools | 45.9% | 10.2% | not listed | 56.9% |
| Humanity's Last Exam, with tools | 57.4% | 18.7% | not listed | 64.5% |
| FrontierCode 1.1 (Main) | 46.4% | not listed | 42.4% | 52.1% (Xhigh) |
The gap between generations is larger than the gap between tiers: Haiku 5.5 beats Haiku 4.5 by far more than Sonnet 5.5 beats Haiku 5.5 on most rows.
Metir analysis of Anthropic's published table
Reading across the table, the generational jump is the larger story: on OSWorld 2.1 the new Haiku moves from 15.7% to 72.4%. The tier gap to Sonnet 5.5 is real but uneven. It is about 11 points on OSWorld and Humanity's Last Exam, but 31 points on Terminal-Bench 4.0, which tests multi-step command-line work. For long, branching terminal tasks, the larger model keeps a clear lead. Against GPT-6 Luna, which shares Haiku 5.5's $0.10 and $0.50 list price according to our GPT-6.1 Sol coverage, Anthropic's table shows Haiku 5.5 ahead on every row where Luna has a score, though a vendor comparing against a rival's model deserves a second look from independent evaluators.

Where a small model fits in a model lineup
The economics suggest a routing pattern rather than a winner. Using the published rates:
- Haiku 5.5 tier: high-volume, latency-sensitive steps such as classification, extraction, routing and subagents, which is the use Anthropic's documentation itself names. At $0.10 input, a million short requests is a line item, not a budget.
- Sonnet 5.5 tier ($2 and $10): the default for agentic work where long-horizon reliability matters, including terminal and coding tasks where the table shows the widest gap. See our Sonnet 5.5 analysis.
- Opus 5.5 tier ($4 and $20): the hardest reasoning and high-stakes steps, with cache reads now at $0.20. See our Opus 5.5 coverage.
The same logic applies across vendors. OpenAI prices GPT-6.1 Sol at $2 and $10 and GPT-6 Luna at $0.10 and $0.50, and Google listed Gemini 3.5 Flash-Lite at $0.30 and $2.50 in our July coverage. By list price, Haiku 5.5 and Luna are level, with Flash-Lite roughly three times higher on input and five times on output. List price is only one input, though. Effort settings, token counts per task, tokenizer differences and rate limits all move the effective cost, so the sound comparison is cost per correct answer on your own workload.
What to watch
- Effort as a cost lever. Higher effort buys accuracy with more thinking tokens. Anthropic's charts show the levels but a team should measure its own accuracy-per-dollar curve before choosing a default.
- The 100K step. A prompt crossing 100,000 tokens costs five times as much per token, so retrieval that trims context may save more than a model swap.
- Independent evaluations. Today's scores come from Anthropic. Third-party leaderboards will show whether the gains hold on other task mixes.
- Cheaper subagents. Cheaper small models make multi-agent designs, with a large planner and many small workers, cheaper to run. Anthropic lists subagent tasks among the intended uses.
Because prices and strengths now shift every few weeks across vendors, many teams keep access to several models and route by task. Metir offers Claude models alongside other providers in one chat, so switching tiers is a menu choice rather than an integration project.
Sources:
- Anthropic: Claude Haiku 5.5 launch page
- Anthropic docs: Claude Haiku 5.5 model overview
- Anthropic docs: Pricing
- Metir: Claude Sonnet 5.5 release
- Metir: GPT-6.1 Sol release and pricing
- Metir: Gemini 3.6 Flash and 3.5 Flash-Lite
Image credits
- Hero: "Dario Amodei at TechCrunch Disrupt 2023 03.jpg" by TechCrunch, licensed CC BY 2.0, via Wikimedia Commons.
- In-article: "Dario Amodei at TechCrunch Disrupt 2023 05.jpg" by TechCrunch, licensed CC BY 2.0, via Wikimedia Commons.
