Anthropic logo

Claude Sonnet 5

Multimodal
Anthropic

Claude Sonnet 5 is Anthropic's most agentic Sonnet-class model, an upgrade to Sonnet 4.6 that narrows the gap to Opus 4.8 on reasoning, tool use, coding, computer use, and knowledge work while staying lower priced. It plans, uses tools like browsers and terminals, and runs autonomously for long-horizon tasks.

Key Specifications

Parameters
-
Context
1.0M
Release Date
June 30, 2026
Average Score
60.5%

Timeline

Key dates in the model's history
Announcement
June 30, 2026
Last Update
August 27, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$2.00
Output (per 1M tokens)
$10.00
Max Input Tokens
1.0M
Max Output Tokens
64.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
Agentic coding (SWE-bench Verified, 500-problem subset). Standard configuration, adaptive thinking at max effort, thinking blocks included in sampling, averaged over 5 trials.Self-reported
85.2%

Other Tests

Specialized benchmarks
ArXivMath
Research-level final-answer math (MathArena, April + May 2026 releases, 81 problems). With tools: 72.2%. Without tools: 65.7%. Extended thinking, averaged over 4 attempts per problem.Self-reported
72.2%
AutomationBench
Zapier end-to-end business-workflow benchmark, private held-out leaderboard set. Max effort.Self-reported
13.5%
BenchCAD
BenchCAD Vision2Code voxel IoU. With Python tools: 0.373. Without tools: 0.266. Random 1,000-file subset averaged over 5 runs.Self-reported
37.3%
BrowseComp
Agentic search. Single-agent with web search, web fetch, code execution, adaptive thinking at max effort, 10M-token limit, context compaction at 200k tokens. Multi-agent reaches 86.6%.Self-reported
84.7%
ChartMuseum
Chart QA requiring visual reasoning (1,162 questions). With Python tools: 86.7%. Without tools: 70.1%. Test split averaged over 5 runs, Claude Sonnet 4.6 grader.Self-reported
86.7%
CharXiv-R
CharXiv Reasoning (1,000 validation questions). With Python tools: 88.3%. Without tools: 77.0%. Averaged over 5 runs, Claude Sonnet 4.6 grader.Self-reported
88.3%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 54% ± 4%.Self-reported
54.0%
FrontierCode
FrontierCode v1, Cognition's production-quality coding evaluation, at max effort.Self-reported
38.8%
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at xhigh effort: 42.7%.Self-reported
42.7%
GDP.pdf
Expert multimodal document QA (Surge AI, 100 prompts/PDFs). Mean criteria pass rate with Python tools: 81.6%. Without tools: 67.5%. Internal harness, averaged over 5 runs.Self-reported
81.6%
HealthBench Professional
Length-adjusted score with adaptive thinking at max effort. Claude Opus 4.8 grader, averaged over 5 trials.Self-reported
57.8%
Humanity's Last Exam
Multidisciplinary reasoning (2,500 questions). With tools (web search, web fetch, code execution): 57.4%. Without tools: 43.2%. Auto thinking, 1M total-token cap, Claude Opus 4.6 grader.Self-reported
57.4%
Legal Agent Benchmark
All-pass rate on Harvey's held-out set (mean criterion-pass rate 91.2%). Full public set all-pass rate: 8.9% (88.26% mean criterion-pass). Adaptive thinking at max effort.Self-reported
5.8%
OfficeQA Pro
Grounded reasoning over U.S. Treasury Bulletin documents (133-question subset). Internal agentic harness with extracted-text documents and code-execution tools. Exact-match, mean of 5 trials. OfficeQA full set: 73.3%.Self-reported
59.4%
OSWorld-Verified
Agentic computer use. 361 tasks at 100 steps, 1080p resolution, adaptive thinking at max effort. First-attempt success rate averaged over 5 runs. Updated harness (zoom-tool bug fix, 128K max tokens per turn).Self-reported
81.2%
SWE-bench Multilingual
300 problems across 9 programming languages. Averaged over 5 trials at max effort.Self-reported
78.3%
SWE-Bench Multimodal
Visual context (screenshots, design mockups) added to issue descriptions. Internal harness, averaged over 5 trials.Self-reported
28.1%
SWE-Bench Pro
Harder SWE-bench variant: actively-maintained repos with larger multi-file diffs and reduced public ground-truth leakage. Averaged over 5 trials at max effort.Self-reported
63.2%
Terminal-Bench 2.0
Agentic terminal coding (Terminal-Bench 2.1) with the mini-SWE-agent harness at xhigh effort. Mean reward over 5 attempts across 89 unique tasks (445 trials).Self-reported
80.4%
Toolathlon
Pass@1 averaged over 3 trials across all 108 tasks. Internal harness with adaptive thinking at max effort. Pass@3: 63.0%, Pass^3: 40.7%.Self-reported
54.3%
USAMO 2026
Proof-based math (USAMO 2026, held March 21-22 2026). MathArena grading methodology (proofs rewritten by a neutral model and judged by a 3-model panel; minimum score taken). High effort, 300k token limit, averaged over 10 attempts. 79.5%.Self-reported
79.5%

License & Metadata

License
proprietary
Announcement Date
June 30, 2026
Last Updated
August 27, 2026

Compare Claude Sonnet 5

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.