Anthropic released Claude Haiku 5.5 on 7 October, cutting per-token prices by up to 90% on its standard-prompt tier and halving Sonnet 5.5’s cache-read cost on the same day. The new small model is Anthropic’s fastest and cheapest yet, built for high-volume narrowly-scoped work — summarisation, classification, database queries and live customer support — and pitched as a worker inside larger agent runs rather than a replacement for Sonnet or Opus. It is live now across AWS, Google Cloud and Microsoft Azure, per Yahoo Finance’s coverage of the release.
90%cheaper per input token than Haiku 4.5 on Anthropic’s standard-prompt tier — a category that covers roughly 90% of all previous Haiku traffic
A clean sweep in the benchmarks
The Decoder’s reporting on Anthropic’s launch benchmarks put Haiku 5.5 ahead of OpenAI’s budget GPT-6 Luna on every test it ran:
- Terminal-Bench 4.0 (coding agents): 39.2% against GPT-6 Luna’s 16.4% and Haiku 4.5’s 0%
- OSWorld 2.1 (computer use — where the model drives a computer on its own): 72.4% against 48.9% and 15.7%
- GDPval-AA v2.1 (knowledge work): 1,620 against 1,437 and 735 — more than double the previous Haiku
- Humanity’s Last Exam with tools: 57.4%, up from 18.7% on Haiku 4.5
Sonnet 5.5 still tops the harder entries — 70.6% on Terminal-Bench 4.0, 83.9% on OSWorld 2.1 — and Anthropic’s positioning is consistent: small model for narrow work and as a worker inside bigger runs; larger models for complex coding. The 5.5 launch completes a family that began with Opus 5.5 on 22 September, which Anthropic framed as the first release in a series where Sonnet and Haiku would follow with the same performance, efficiency and safety improvements.
Yahoo Finance notes one early customer — workflow platform Asana — reported a 30%+ drop in task-completion latency and up to 2.5x faster inference after switching. That is the supporting evidence behind the “cheap and fast” pitch — the kind of number a UK operations lead wants before retraining a support pipeline on the new tier.
The saving comes with a footnote
The 90% headline is genuine but the realised saving is lower. Haiku 5.5 ships with an updated way of counting inputs and outputs, and the new counter consumes slightly more per task than Haiku 4.5 — the same pattern The Decoder noted with the Opus 4.x line, where per-task use jumped about 30% from the counting change alone. For UK teams running customer-support automation inside the standard-prompt range the headline saving holds; for big-document agents above that ceiling, the gains are more modest. Prompts longer than the standard tier cost five times the headline rate, per The Decoder’s reporting.
We’ve covered this cheaper-model, bigger-bill trap before; the rule of thumb is the same here: weigh the headline cut against your own request counts, not the per-token number in isolation.
The bigger win for coding-heavy teams sits with Sonnet 5.5. Cache reads — the part of the bill where the API charges much less for prompts you’ve already sent that day — drop from $0.20 to $0.10 per million tokens, an automatic change The Decoder says shaves around 20% off typical coding-agent runs. For a UK development team whose Claude Code bill runs heavy on cache reuse, that half-decimal change may be the day’s larger win. Monthly API credits of $100–$500 for Max and Team subscribers add a way to experiment with both models without opening a procurement line.
What to do with this
Three things are worth doing this week, in roughly the order they matter:
- Re-price existing Haiku 4.5 workloads. If you have ticket triage, support email classification or structured-extraction jobs running on Haiku 4.5 at standard prompt sizes, the cost ceiling has dropped by nearly an order of magnitude. What was a pilots-only line on a vendor bill becomes viable for production.
- Pull the cache-read line off your Sonnet 5.5 invoice. Long Claude Code runs — codebase audits, agent passes, refactors — produce a cache-read figure that is now half what it was. Surface the line for finance so the saving lands on this month’s run rather than waiting on a vendor statement.
- Try the adjustable reasoning knob. Haiku 5.5 is the first small Claude with a tunable reasoning depth. Sub-agent calls where the parent agent already knows the answer is narrow — picking a canned reply from context, filling a JSON template from a record — can dial reasoning down to cut output cost further without harming quality.
The wider picture: running a Claude small model in volume has been a pilots-only cost for most UK small firms until now. At these prices it is a viable option for workloads that previously had to live on an open-weight self-host — though self-hosting still wins on data residency for regulated loads. Anthropic’s timing reads as a competitive shot at OpenAI’s budget tier, and a reminder that the agent-cost story for next year’s budget has moved again.
Sources & quotes
Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →


