Anthropic logo

Claude Sonnet 5.5

Multimodal
Anthropic

Claude Sonnet 5.5 is the second model in Anthropic's Claude 5.5 family and a faster, lower-cost complement to Claude Opus 5.5. It targets well-scoped everyday tasks, coding, and polished documents, slides, and spreadsheets, and runs 30%+ faster than Claude Sonnet 5 while typically using far fewer tokens at the same $2/$10 per million input/output prices. It accepts text and image input with text output, supports multilingual and vision workloads, tool use, and adaptive thinking (Claude Platform default effort high; Claude Code and apps default medium). Context window is 1M tokens with 128K max output on the sync Messages API; reliable knowledge and training-data cutoffs are June 2026. Cache reads are $0.20 and 5-minute cache writes are $2.50 per million tokens. Capability scores are self-reported from the Claude Sonnet 5.5 System Card (Table 8.1.A and §8) and the launch announcement. Available on the Claude API as `claude-sonnet-5-5`.

Key Specifications

Parameters
-
Context
1.0M
Release Date
September 28, 2026
Average Score
67.4%

Timeline

Key dates in the model's history
Announcement
September 28, 2026
Last Update
September 29, 2026
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
June 1, 2026
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.00
Output (per 1M tokens)
$10.00
Max Input Tokens
1.0M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
AA-Briefcase v1.1
Announcement + System Card Table 8.1.A. AA-Briefcase v1.1 Elo (max_score 3000). Same Artificial Analysis pre-release structured-output bug caveat as GDPval-AA. • Self-reported
60.4%
ArXivMath
System Card §8.9. ArXivMath August 2026 (57 problems); without tools; max effort; averaged over four attempts per problem. Evaluated internally without safeguards classifiers. • Self-reported
86.8%
ArXivMath (with tools)
System Card §8.9. ArXivMath August 2026 (57 problems); with tools (code-execution sandbox, no internet); max effort; averaged over four attempts per problem. • Self-reported
95.2%
AutomationBench
System Card Table 8.1.A / §8.14.6. Zapier AutomationBench 44.7% at max effort with API default fallbacks enabled. Adaptive thinking at max effort. • Self-reported
44.7%
BenchCAD
System Card §8.13.2. BenchCAD Vision2Code 1k subset; voxel IoU 0.747 without tools; adaptive thinking at max effort; averaged over five runs. • Self-reported
74.7%
BenchCAD (with Python tool)
System Card §8.13.2. BenchCAD Vision2Code 1k subset; voxel IoU 0.963 with tools; adaptive thinking at max effort; averaged over five runs. • Self-reported
96.3%
Biomedical image analysis
System Card §8.17.7. Biomedical image analysis 72.2%; 1 attempt per problem; adaptive thinking at max effort. • Self-reported
72.2%
BioMysteryBench
System Card §8.17.1. BioMysteryBench Human Solvable subset 89.2%; adaptive thinking at max effort; bash/file tools with package allow-list. (RSP summary rounded to 0.89; §8.17.1 is authoritative.) • Self-reported
89.2%
BioMysteryBench (Human Difficult)
System Card §8.17.1. BioMysteryBench Human Difficult subset 44.7%; adaptive thinking at max effort; bash/file tools with package allow-list. (RSP summary rounded to 0.45; §8.17.1 is authoritative.) • Self-reported
44.7%
Chartography
System Card §8.13.1. Chartography with tools 90.2%; adaptive thinking at max effort; container with image file, standard libraries, and image cropping tool. • Self-reported
90.2%
Chartography (no tools)
Announcement + System Card §8.13.1. Chartography without tools 61.6%; adaptive thinking at max effort. • Self-reported
61.6%
CursorBench 4.0
Announcement table + System Card §8.8. CursorBench 4.0 at max effort 55.5% (independently measured by Cursor in its production harness). Also reported: xhigh 53.1%, high 47.8%, medium 39.2%. • Self-reported
55.5%
DeepSWE 1.1
System Card §8.3. DeepSWE v1.1; 5-trial average 71.0%. Adaptive thinking at max effort. • Self-reported
71.0%
De novo protein binder design
System Card §8.17.6. De novo protein binder design 82.3%; 15 targets, 30 miniprotein binders each, 24h wall-clock budget; adaptive thinking at max effort. • Self-reported
82.3%
FrontierCode 1.1
Announcement + System Card §8.4 / Table 8.1.A. FrontierCode v1.1 Main. Catalog stores Xhigh 52.1% as the fair headline (Table 8.1.A max-effort Main is 46.2%). Anthropic notes Max is lower due to Claude Code code-review subagent timeouts / out-of-scope edits penalized by mergeability grading. Extended Xhigh 64.4% stored separately. • Self-reported
52.1%
FrontierCode 1.1 (Extended)
System Card §8.4. FrontierCode v1.1 Extended at Xhigh effort 64.4% (max effort Extended 59.1%). Same mergeability grading caveats as Main. • Self-reported
64.4%
FrontierSWE V2
System Card §8.7. FrontierSWE v2; Proximal harness; max effort; 5 trials per task; mean across trials. • Self-reported
61.9%
GDPval-AA 2.1
Announcement + System Card Table 8.1.A. GDPval-AA v2.1 Elo (max_score 3000). Artificial Analysis evaluated a pre-release Claude Platform deployment with a structured-output bug later fixed; Anthropic expects any impact to be small and to understate performance. • Self-reported
61.5%
Global-MMLU
System Card §8.16.1. GMMLU average accuracy 92.1% across 42 languages; adaptive thinking at max effort; single trial; no tools. • Self-reported
92.1%
HealthBench
System Card §8.15.4. HealthBench length-adjusted scores across all five effort levels sit in 64.7–65.4%; store the published band max 65.4% to match the catalog convention (length-adjusted on healthbench). Raw 69.4% at max effort (§8.15.1) is stored separately as healthbench-raw. • Self-reported
65.4%
HealthBench Professional
System Card Table 8.1.A length-adjusted score 69.2% at max effort. Raw (pre-length-adjustment) 77.1% is stored separately as healthbench-professional-raw (§8.15.2). Adaptive thinking at max effort. Length-adjusted per HealthBench Professional paper method. • Self-reported
69.2%
HealthBench Professional (raw)
System Card §8.15.2. HealthBench Professional raw (before length adjustment) 77.1% at max effort; level with Opus 5.5 raw. Length-adjusted 69.2% is stored as healthbench-professional (Table 8.1.A). • Self-reported
77.1%
HealthBench (raw)
System Card §8.15.1. HealthBench raw (before length adjustment) 69.4% at max effort (slightly ahead of Opus 5.5 raw 68.1%). Length-adjusted band max 65.4% is stored as healthbench (§8.15.4). • Self-reported
69.4%
Humanity's Last Exam (no tools, text-only)
System Card Table 8.1.A / §8.11.1. HLE no tools (reasoning-only); adaptive thinking at max effort. • Self-reported
56.9%
Humanity's Last Exam (with tools)
Announcement + System Card Table 8.1.A / §8.11.1. HLE with tools (web search, web fetch, programmatic tool calling, code execution); adaptive thinking at max effort. • Self-reported
64.5%
LatchBio SingleCellBench
System Card §8.17.2. LatchBio SingleCellBench 59.1%; adaptive thinking at max effort; bash/file tools. • Self-reported
59.1%
LatchBio SpatialBench Verified
System Card §8.17.2. LatchBio SpatialBench Verified 72.5%; adaptive thinking at max effort; bash/file tools. • Self-reported
72.5%
Legal Agent Benchmark
System Card §8.14.2. Legal Agent Benchmark held-out set; all-pass rate 11.7% at high effort (AA harness). Mean criterion-pass rate 92.1% noted in card but not stored as a separate catalog metric (all-pass is the standard LAB reporting id). At max effort: 10.0% all-pass / 93.1% mean criterion-pass. • Self-reported
11.7%
Medicinal Chemistry (ADME)
System Card §8.17.4. Medicinal Chemistry ADME 65.3% (504 questions); adaptive thinking at max effort. • Self-reported
65.3%
MILU
System Card §8.16.2. MILU average accuracy 91.6% across 11 languages; adaptive thinking at max effort; averaged over 5 trials; no tools. • Self-reported
91.6%
Morphology-to-molecule matching
System Card §8.17.3. Axiom Bio morphology-to-molecule matching 25.0%; adaptive thinking at max effort. • Self-reported
25.0%
OfficeQA
System Card §8.14.1. OfficeQA full set; max effort; agentic harness with extracted-text documents and code-execution tools. • Self-reported
76.9%
OfficeQA Pro
System Card §8.14.1. OfficeQA Pro (133-question subset); max effort. • Self-reported
65.6%
OSWorld 2.1 (partial)
Announcement + System Card Table 8.1.A / §8.13.3. OSWorld 2.1 partial-credit score 80.1% (Pass@1, 5-run average) at max effort. Not OSWorld 2.0. • Self-reported
80.1%
OSWorld 2.1 (strict)
System Card §8.13.3. OSWorld 2.1 strict pass rate 43.5% (every checkpoint satisfied; Pass@1, 5-run average) at max effort. Not OSWorld 2.0. • Self-reported
43.5%
PhysicianBench
System Card §8.15.3. PhysicianBench 100 physician EHR tasks; 63.2% pass rate at max effort. Safety classifiers enabled; graded with Claude Opus 5. • Self-reported
63.2%
Program Bench
System Card §8.10.1. ProgramBench 166 golden tasks; mini-swe-agent harness; no 6-hour time limit; hidden test pass rate against tests the reference binary passes. • Self-reported
79.7%
Protein Design — Library Ranking
System Card §8.17.5. Protein Design Library Ranking 54.8%; 3 attempts per problem; adaptive thinking at max effort. • Self-reported
54.8%
Protein Design — Sequence Generation
System Card §8.17.5. Protein Design Sequence Generation 51.0%; 1 attempt per problem; adaptive thinking at max effort. • Self-reported
51.0%
Protocols Troubleshooting
System Card §8.17.8. Protocols Troubleshooting 67.3%; adaptive thinking at max effort. • Self-reported
67.3%
Protocols Understanding (V2)
System Card §8.17.8. Protocols Understanding (V2) (Benchling) 66.6%; adaptive thinking at max effort. • Self-reported
66.6%
SWE-bench Multilingual
System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. 300 problems across nine programming languages. • Self-reported
90.3%
SWE-Bench Multimodal
System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. Visual context (screenshots and design mockups). • Self-reported
54.3%
SWE-Bench Pro
System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. • Self-reported
81.3%
Terminal-Bench 4.0
Announcement table + System Card §8.5. Terminal-Bench 4.0 accuracy 70.6% with production safeguards enabled (Claude Code --bare, max thinking effort; 5-trial average). Safeguard-flagged requests used the default server-side fallback (1.2% of requests). • Self-reported
70.6%
Terminal-Bench-Science 0.1
System Card §8.6. Terminal-Bench-Science 0.1; 59.9% with safeguards enabled (no fallbacks fired). Claude Code --bare, max thinking effort. • Self-reported
59.9%
Toolathlon-Verified
System Card §8.14.5. Toolathlon Verified Pass@1 77.8% (3-trial average across 108 tasks). Adaptive thinking at max effort; safety/sandbox stops counted as failures (0.6% of trials). • Self-reported
77.8%

License & Metadata

License
proprietary
Announcement Date
September 28, 2026
Last Updated
September 29, 2026

Compare Claude Sonnet 5.5

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.