Claude Sonnet 5
MultimodalClaude Sonnet 5 is Anthropic's most agentic Sonnet-class model, an upgrade to Sonnet 4.6 that narrows the gap to Opus 4.8 on reasoning, tool use, coding, computer use, and knowledge work while staying lower priced. It plans, uses tools like browsers and terminals, and runs autonomously for long-horizon tasks.
Key Specifications
Parameters
-
Context
1.0M
Release Date
June 30, 2026
Average Score
60.5%
Timeline
Key dates in the model's history
Announcement
June 30, 2026
Last Update
August 27, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$2.00
Output (per 1M tokens)
$10.00
Max Input Tokens
1.0M
Max Output Tokens
64.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
Agentic coding (SWE-bench Verified, 500-problem subset). Standard configuration, adaptive thinking at max effort, thinking blocks included in sampling, averaged over 5 trials. • Self-reported
Other Tests
Specialized benchmarks
ArXivMath
Research-level final-answer math (MathArena, April + May 2026 releases, 81 problems). With tools: 72.2%. Without tools: 65.7%. Extended thinking, averaged over 4 attempts per problem. • Self-reported
AutomationBench
Zapier end-to-end business-workflow benchmark, private held-out leaderboard set. Max effort. • Self-reported
BenchCAD
BenchCAD Vision2Code voxel IoU. With Python tools: 0.373. Without tools: 0.266. Random 1,000-file subset averaged over 5 runs. • Self-reported
BrowseComp
Agentic search. Single-agent with web search, web fetch, code execution, adaptive thinking at max effort, 10M-token limit, context compaction at 200k tokens. Multi-agent reaches 86.6%. • Self-reported
ChartMuseum
Chart QA requiring visual reasoning (1,162 questions). With Python tools: 86.7%. Without tools: 70.1%. Test split averaged over 5 runs, Claude Sonnet 4.6 grader. • Self-reported
CharXiv-R
CharXiv Reasoning (1,000 validation questions). With Python tools: 88.3%. Without tools: 77.0%. Averaged over 5 runs, Claude Sonnet 4.6 grader. • Self-reported
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 54% ± 4%. • Self-reported
FrontierCode
FrontierCode v1, Cognition's production-quality coding evaluation, at max effort. • Self-reported
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at xhigh effort: 42.7%. • Self-reported
GDP.pdf
Expert multimodal document QA (Surge AI, 100 prompts/PDFs). Mean criteria pass rate with Python tools: 81.6%. Without tools: 67.5%. Internal harness, averaged over 5 runs. • Self-reported
HealthBench Professional
Length-adjusted score with adaptive thinking at max effort. Claude Opus 4.8 grader, averaged over 5 trials. • Self-reported
Humanity's Last Exam
Multidisciplinary reasoning (2,500 questions). With tools (web search, web fetch, code execution): 57.4%. Without tools: 43.2%. Auto thinking, 1M total-token cap, Claude Opus 4.6 grader. • Self-reported
Legal Agent Benchmark
All-pass rate on Harvey's held-out set (mean criterion-pass rate 91.2%). Full public set all-pass rate: 8.9% (88.26% mean criterion-pass). Adaptive thinking at max effort. • Self-reported
OfficeQA Pro
Grounded reasoning over U.S. Treasury Bulletin documents (133-question subset). Internal agentic harness with extracted-text documents and code-execution tools. Exact-match, mean of 5 trials. OfficeQA full set: 73.3%. • Self-reported
OSWorld-Verified
Agentic computer use. 361 tasks at 100 steps, 1080p resolution, adaptive thinking at max effort. First-attempt success rate averaged over 5 runs. Updated harness (zoom-tool bug fix, 128K max tokens per turn). • Self-reported
SWE-bench Multilingual
300 problems across 9 programming languages. Averaged over 5 trials at max effort. • Self-reported
SWE-Bench Multimodal
Visual context (screenshots, design mockups) added to issue descriptions. Internal harness, averaged over 5 trials. • Self-reported
SWE-Bench Pro
Harder SWE-bench variant: actively-maintained repos with larger multi-file diffs and reduced public ground-truth leakage. Averaged over 5 trials at max effort. • Self-reported
Terminal-Bench 2.0
Agentic terminal coding (Terminal-Bench 2.1) with the mini-SWE-agent harness at xhigh effort. Mean reward over 5 attempts across 89 unique tasks (445 trials). • Self-reported
Toolathlon
Pass@1 averaged over 3 trials across all 108 tasks. Internal harness with adaptive thinking at max effort. Pass@3: 63.0%, Pass^3: 40.7%. • Self-reported
USAMO 2026
Proof-based math (USAMO 2026, held March 21-22 2026). MathArena grading methodology (proofs rewritten by a neutral model and judged by a 3-model panel; minimum score taken). High effort, 300k token limit, averaged over 10 attempts. 79.5%. • Self-reported
License & Metadata
License
proprietary
Announcement Date
June 30, 2026
Last Updated
August 27, 2026
Compare Claude Sonnet 5
All comparisonsSimilar Models
All ModelsClaude Sonnet 4
Anthropic
MM
Best score:0.8 (GPQA)
Released:May 2025
Price:$3.00/1M tokens
Claude Opus 4.7
Anthropic
MM
Best score:0.9 (GPQA)
Released:Apr 2026
Price:$5.00/1M tokens
Claude Opus 4
Anthropic
MM
Best score:0.8 (GPQA)
Released:May 2025
Price:$15.00/1M tokens
Claude Fable 5
Anthropic
MM
Released:Jun 2026
Price:$10.00/1M tokens
Claude Opus 5
Anthropic
MM
Released:Jul 2026
Price:$5.00/1M tokens
Claude 3.5 Sonnet
Anthropic
MM
Best score:0.9 (HumanEval)
Released:Jun 2024
Price:$3.00/1M tokens
Claude 3 Opus
Anthropic
MM
Best score:1.0 (ARC)
Released:Feb 2024
Price:$15.00/1M tokens
Claude Opus 4.6
Anthropic
MM
Best score:1.0 (TAU)
Released:Feb 2026
Price:$5.00/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.