All News
anthropicclaudemodel-releasebenchmarksai-pricingai-safety

Claude Opus 5.5 ships 20% cheaper, on a benchmark table that scores a model you cannot call

Anthropic's Claude Opus 5.5 is 20% cheaper than Opus 5 and leads its own benchmark table, though a footnote says safeguards blunted the biology and cyber rows.

Vlad MakarovVlad Makarovreviewed and published
6 min read
Claude Opus 5.5 ships 20% cheaper, on a benchmark table that scores a model you cannot call

Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in a new Claude 5.5 family and its first launch since chief executive Dario Amodei published an essay arguing the frontier should be paced, not hurried. Its own announcement states the claim it rests on: the model "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." The cost half is checkable arithmetic. The capability half rests on a scoreboard Anthropic built and, as its footnotes show, partly stepped around.

The price cut is the part you can check

UnitOpus 5.5Opus 5Change
Input, per 1M tokens$4$5-20%
Output, per 1M tokens$20$25-20%
Cache reads, per 1M tokens$0.20$0.50-60%
Cache writes, per 1M tokens$5$6.25-20%
Fast mode, in / out$8 / $40not offeredup to 2.5x speed
Batch pricing50% off50% offunchanged

Anthropic attributes the 40% to how many tokens the model spends reaching an answer, not to those prices, and says output arrives over 30% faster than Opus 5's. Neither is reproducible from the launch material. What a buyer can act on is narrower: cheaper tokens, cheaper cache reads, a fast tier at double the standard rate, the same batch discount. Usage limits rose on Pro, Max, Team and seat-based Enterprise plans, and subscribers get a rate-limit reset they can hold.

Reading Anthropic's own scoreboard

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.066.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 Main54.4%50.3%48.0%53.3%47.5%
CursorBench 4.057.8%51.8%46.6%41.7%
GDPval-AA v2.11,8461,7351,7081,5421,588
AutomationBench40.0%31.4%26.9%41.4%28.8%
Humanity's Last Exam, with tools67.7%65.6%63.6%57.2%
Terminal-Bench-Science 0.158.7%52.6%29.0%64.6%22.4%
OSWorld 2.081.8%80.7%74.0%
Chartography, with tools89.0%88.4%83.4%

Every figure except OpenAI's is Anthropic's own measurement on its harnesses; the OpenAI columns are "as reported by OpenAI," a rival's numbers quoted, not re-run. Dashes mark empty cells; Anthropic's notes call the OSWorld and Chartography rows partial. Opus 5.5 tops seven of nine rows, with 10.6 points over Fable 5.1 on Terminal-Bench 4.0, 6 on CursorBench, 8.6 on AutomationBench, which Zapier ran. Two rows run the other way — GPT-6 Astra takes Terminal-Bench-Science, 64.6% to 58.7%, and edges AutomationBench, 41.4% to 40.0%, results Anthropic printed in its own table.

What the footnotes concede

Anthropic concedes the awkward reading of its own numbers. "At these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences," the post reads, adding that "the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest."

A footnote narrows it further. Because Opus 5.5 sits close to Mythos 5.1 in biology and cybersecurity, safeguard interventions meant cyber tasks in the suite were completed by Opus 4.8 and biology and frontier-LLM-development tasks by Opus 5, "which likely reduces Claude Opus 5.5's performance on these benchmarks." The model is scored on a version of itself customers never call — a defensible safety choice and a real problem for anyone reading the table as a ranking.

Safety, shipped as a routing decision

Anthropic calls this its best score yet on its automated behavioral audit and says the model resists prompt injection better than Opus 5. Reuters, which confirmed the company's account of pre-release testing by Frontier Design and METR, reports Opus 5.5 was about 85% less likely than Opus 5 or Mythos 5.1 to attempt to bypass containment boundaries in a dedicated evaluation.

Capability brings the constraint. Most cybersecurity requests are re-routed to Opus 4.8. Vetted organizations can apply to the Life Sciences Verification Program today, and the Cyber Verification Program expands to Opus 5.5 "in the coming weeks" with three tiers of trusted access. Thinking can no longer be switched off, and "preserved thinking," framed as anti-distillation, applies to API accounts created on or after August 31, 2026. The EU AI Act's watermarking rules apply too.

Four breaking changes, one quiet one

The platform documentation lists the model as claude-opus-5-5 on the Claude API, AWS Bedrock, Google Cloud and Microsoft Foundry, with a 1M-token context window, 128K maximum output (300K with a batch beta header), adaptive thinking always on at a default effort of "medium," knowledge cutoff June 2026, and retirement no sooner than September 22, 2027.

Four changes break code that worked against Opus 5: thinking can no longer be disabled, forced tool use returns an error, thinking blocks are tied to the model and conversation that produced them, and the older computer_20251124 computer-use tool is not accepted on the API or Google Cloud. The quiet one is worse for streaming interfaces: text between tool calls now arrives inside thinking blocks that are empty by default, so an app rendering that stream as progress goes quiet until its authors set a display value.

The pacing question, answered sideways

Reuters framed the launch as landing amid debate over AI risks, days after Amodei urged rivals to slow the pace of releasing new capabilities — the argument this site covered when competitors and Congress answered it. Anthropic's answer is that Opus 5.5 does frontier work more cheaply — a "performs at the level of Fable 5.1" release, not a capability jump, which is different shipping than the pacing essay warned about. Whether that distinction survives contact with buyers is open; the Fable 5.1 and Mythos 5.1 release shows how fast the company's own framing gets read as a scoreboard. Sonnet 5.5 and Haiku 5.5 arrive "in the coming weeks," and OpenAI cut GPT-6 prices the same evening.

Would an audit settle it?

Three things, none exotic. Someone with access could re-run the benchmark suite behind the table — including the biology and cyber tasks the safeguards diverted — so the scores describe the model customers actually get. The token-efficiency claim under the 40% could be measured on a neutral harness at matched effort, the load-bearing number for the cost story, with no published method. And the safeguard routing could be published as a rule, not a footnote, so a developer knows which requests their Opus 5.5 call will not answer. The system card documents the containment-boundary evaluation, not the benchmark table. Until someone outside Anthropic does, "at the level of Fable 5.1 at 40% less" is a price point with a plausible capability claim attached, and only one half is checkable today.

Related Articles

Scroll down

to load the next article