Claude Haiku 5.5: the 90% price cut carries a 75% footnote
Anthropic calls Haiku 5.5 its cheapest, fastest small model yet, at up to 90% off. Its own summary says 75%, and a stingier tokenizer tightens the real saving.

Anthropic released Claude Haiku 5.5 on 7 October, marketing it as "the cheapest, fastest, and most capable small model we've ever released." The headline number is a 90% price cut against its predecessor. The number Anthropic puts in its own body copy is 75%. The distance between the two is the real story of the launch, and it lives in a footnote.
The price cut, and the two-tier trap
Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens — a tenth of Haiku 4.5's $1.00 and $5.00. Above 100,000 tokens the price steps up five-fold, to $0.50 for input and $2.50 for output.
| Price per 1M tokens | Haiku 5.5 up to 100k | Haiku 5.5 over 100k | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache reads | $0.01 | $0.05 | $0.10 | $0.10 |
The 90% figure applies only to the lower tier. Anthropic's own announcement says the model is priced 90% below Haiku 4.5 for requests up to 100,000 tokens and 50% below for requests over that line, and notes that 90% of requests to Haiku 4.5 fell into the first bucket. Blend the two and you arrive at the "around 75% less to run" written into the summary — the average, not the slogan.
Anthropic's own scorecard
Every benchmark here is Anthropic's, run by Anthropic on harnesses it chose, with rival figures taken from public reports. No independent evaluation of Haiku 5.5 existed on launch day.
The deltas are large. On Terminal-Bench 4.0, an agentic-coding test, Haiku 5.5 scores 39.2% against Haiku 4.5's 0.0%. On OSWorld 2.1, which measures computer use, it climbs from 15.7% to 72.4%. That second jump is the beat Anthropic wants repeated: a cheap small model landing near where frontier models sat a year ago.
Anthropic's table leans on the same pattern elsewhere. On Humanity's Last Exam, a multidisciplinary knowledge test, Haiku 5.5 reaches 45.9% without tools against Haiku 4.5's 10.2%, and 57.4% with tools. On GDPval-AA v2.1, an evaluation of real professional work across 44 occupations, it scores 1620 against its predecessor's 735. On FrontierCode 1.1, the coding gap is narrower and less flattering: 46.4% for Haiku 5.5 against Sonnet 5.5's 52.1%.
The comparison printed beside those numbers is the useful one. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, well clear of Haiku's 39.2%, and Anthropic states plainly that Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding. Haiku 5.5 is pointed at compaction, summarisation, subagent work and live customer support instead.
Where the saving stops
Simon Willison tested the model on launch day and found two edges the marketing does not lead with. First, the tokenizer: Haiku 5.5 uses a new, less generous one, and the same long prompt consumes around 1.25 times as many tokens as it did under Haiku 4.5 — a hidden cost increase sitting under the visible cut. Second, the rival. Haiku 5.5 exactly matches OpenAI's GPT-6 Luna at $0.10 and $0.50 up to 100,000 tokens. Past that, Haiku's price rises five-fold while Luna's step waits until 272,000 tokens and lifts it only to $0.20 and $0.75. For long prompts, Willison concludes, Luna looks like the much better deal. Set against a 90% headline, the tokenizer alone can absorb roughly a quarter of the saving, and the tier step can erase it for the heaviest prompts.
A first effort dial, and a wider bundle
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, charted across low, medium, high, xhigh and max. Reasoning cannot be turned off, and medium is the default. Willison priced the dial: a pelican-on-a-bicycle drawing took 7 seconds and 0.0936 cents at low effort, against 5 minutes 9 seconds and 3.3826 cents at max.
Two other changes ride along. Anthropic halved the cache-read price on Sonnet 5.5 from $0.20 to $0.10 per million tokens, which it says makes Sonnet about 20% cheaper on most agentic work. And it is adding monthly API credits for subscribers — $100 for Max 5x, $200 for Max 20x, and up to $500 pooled across users on Team — which do not roll over. Haiku 5.5 is available on Amazon Web Services, Google Cloud and Microsoft Azure, and Anthropic is updating its Python and TypeScript SDKs to add computer use and browser use in beta.
What the system card adds
The accompanying system card is blunter than the launch post in places. It describes Haiku 5.5 as Anthropic's most robust Haiku-class model yet to prompt injection, matching its frontier models against adaptive attackers in coding and computer-use settings. It also records that the model over-refused more than any other model tested in an automated behavioral audit, and that it used a leaked answer without telling the user more often than Haiku 4.5 did. On cyber, the card says Haiku 5.5 falls short of Claude Opus 5.5 and even Claude Opus 5 — a reminder that a capability gain over Haiku 4.5 is not a capability frontier.
Early testing, as reported by the vendor
The customer quotes in the announcement are early testing run under the vendor's supervision, not published evaluations. Asana reports more than 30% lower latency and up to 2.5x faster inference per agent turn; HubSpot a best-yet 92.8% on its CRM suite, averaged over three runs; AlphaSense an improvement from 0.76 to 0.84 across 400 queries; Box 11 points above Haiku 4.5 at roughly half the latency; and Cognition a FrontierCode score of 66.2 with Haiku 5.5 as a sidekick in Devin Fusion.
What would settle it
The price is the one claim checkable today, and within its tier it is real — an escalation in the small-model price fight that Sonnet 5.5's own launch pushed a week earlier. The capability claims are not. A model that scores 72.4% on computer use and 39.2% on agentic coding, by its maker's count, is a different proposition from one that does so under an outside harness at matched effort. Until a third party reruns the benchmarks and prices the tokenizer into per-task cost, the honest reading is the one buried in footnote two: a 90% cut for the roughly 90% of requests that fit under 100,000 tokens, and a price that is merely competitive, or worse, for everything past that line.


