Claude Sonnet 5.5 is real: the registry string became a product, and the scorecard is Anthropic's own
Anthropic shipped Claude Sonnet 5.5 two days after a leak named it. The registry string became a product, and the scorecard it arrived with is vendor-run.

Anthropic published Claude Sonnet 5.5 on September 28, two days after the slug appeared in a third-party client's registry with no official page behind it. The registry string became a product: a model id, a price, a model card, a benchmark table. Almost nothing else from the leak survived the trip.
The launch frames Sonnet 5.5 as "the second model in the Claude 5.5 family," a faster, lower-cost complement to Claude Opus 5.5 rather than a challenger to it — strongest at well-scoped everyday work, bug fixes, and polished documents, slides and spreadsheets. It runs 30%+ faster than Sonnet 5 and costs up to 30% less for most work, and Claude Haiku 5.5 "will join the Claude 5.5 family in the coming weeks."
The September 27 leak was precise about a slug and a feature flag and vague about everything else. Its headline detail, an 872K context window with 128K output, was a registry frame inside another company's app. Two days later the launch post advertises no context window at all.
The scorecard is entirely vendor-reported
Everything below is vendor measurement: Anthropic's launch post, with the Sonnet 5, Opus 5.5 and GPT-6 Sol columns taken from those vendors' own published reports, not a shared harness.
| Benchmark | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1, main set | 46.2% (Max), 52.1% (Xhigh) | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | 1487 |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | 1483 |
| Humanity's Last Exam, with tools | 64.5% | 54.9% | 67.7% | — |
| OSWorld 2.1, computer use | 80.1% | 57.0% | 81.8% | — |
| Chartography, no tools | 61.6% | 15.6% | 64.4% | 53.6% |
Terminal-Bench 4.0 does the heavy lifting: 70.6% against Sonnet 5's 10.3%, a jump that only means something if someone outside Anthropic reproduces it. It is also the row least likely to get that treatment.
Anthropic's second footnote has Sonnet 5.5 scoring lower at Max effort than at Xhigh on FrontierCode, because at Max it more often ran Claude Code's code-review skill, which splits review across subagents. In two cases Cognition examined, that produced a timeout or edits beyond the task's scope, and a lower score. The cause is a workflow failure mode — an agent subcontracting its own review — not a scoring quirk.
Two footnotes muddy the rival column from both directions. OpenAI recently fixed a bug that degraded image understanding in GPT-6 Sol, so the AA-Briefcase, GDPval-AA and Chartography figures may not reflect the current build. Anthropic's own numbers for those two came from a pre-release deployment Artificial Analysis ran with a structured-output bug, which Anthropic expects understated its performance. Neither side of the rival column is clean, and Anthropic's own hedge is worth keeping: benchmark scores "capture only one facet of a model's capabilities," and Opus 5.5 "remains clearly stronger at complex, open-ended work requiring sustained judgment."
Same price, and a per-task claim with a regime attached
Pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output tokens, $0.20 per million cache reads. On top of that Anthropic claims Sonnet 5.5 "typically needs far fewer tokens to do the same work," and that in its testing it "costs up to 30% less per task than its predecessor." Read that as a regime, not a list price: per-task cost depends on the effort setting and how much of the prompt lands as a cache read, so "up to 30% less" assumes Anthropic's default effort and its own cache-hit profile.
The independent counterweight is sharper. Artificial Analysis, which ran the model at all five effort settings, measured the opposite at the top: at Max effort it counted roughly 193,000 output tokens per task, "the highest token use we have measured," and priced the run at $7.60 per task, about 50% above Sonnet 5. Both can be true: at Low and Medium effort Anthropic's own testing has Sonnet 5.5 beating Sonnet 5's best score for around a tenth of the per-task cost. Same weights, different regime.
Why a mid-tier model gets frontier safeguards
Sonnet 5.5 is the first Sonnet to ship with cyber safeguards and fallbacks like those Anthropic built for its most capable models, because its cybersecurity capabilities are comparable to Opus 5's; its biology safeguards are unchanged from Sonnet 5's. It is also the first Sonnet with classifiers that prevent reasoning extraction, the defence against accounts harvesting capabilities at scale, and it expands preserved thinking so Claude's thinking cannot be decoupled from the account that created it. Developers who move conversations between accounts, including switching accounts mid-session in Claude Code, are told to read the docs.
The reason a mid-tier model needs that apparatus is in the scorecard: a model within a couple of points of Opus 5.5 on agentic coding, sold at a fraction of the output price, is exactly what distillation attacks target. The system card reports that Sonnet 5.5 crosses no new Responsible Scaling Policy thresholds and stays less capable than Opus 5.5 overall, while noting it develops sophisticated exploits far more capably than Sonnet 5.
Availability, migration, and an unmentioned context window
Sonnet 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud and Microsoft Azure, under the model id claude-sonnet-5-5, with zero data retention. Developers who run Sonnet with thinking off must move to the new between_tools setting before switching. Reuters placed the launch inside Anthropic's buildout ahead of a planned IPO.
The leak could not have shown any of this: pricing, benchmarks, safeguards and availability were all unknowable from a feature flag. On context, Artificial Analysis independently reports a 1 million token window, unchanged from Sonnet 5 — which, if it holds, is not the number the registry string promised. The Opus 5.5 launch still defines the top of the family.
What would settle it
Three things would move Sonnet 5.5 from a press release to a measured product. None exists yet. A third-party cost-per-task figure at a controlled effort setting, since vendor and independent numbers diverge across settings. An independent cyber evaluation, because "first Sonnet with frontier-grade safeguards" is a claim about capability thresholds only an outside lab can score. And a token-efficiency measure at a realistic cache-hit rate, where the 30% saving either survives or evaporates.
Also unreleased: the training corpus beyond the system card's proprietary-mix description, the supervised fine-tuning and reinforcement-learning recipe, and any independent evaluation at launch. The Hacker News submission drew roughly 866 points and 600 comments across September 28 and 29, much of it on the two questions the post cannot answer: token use per task, and whether the price holds at volume. The post's own hedge is the honest summary — the fastest, cheapest Sonnet yet, and not the model Opus 5.5 buyers were already paying for.


