Anthropic logo

Claude Sonnet 4.6

Multimodal
Anthropic

Claude Sonnet 4.6 is a complete upgrade to the Sonnet-class model with improvements in coding, computer use, long-context reasoning, agent planning, knowledge work, and design. Users preferred Sonnet 4.6 over Sonnet 4.5 approximately 70% of the time. The first Sonnet-class model with a 1M token context window (beta) and context compaction. Significant improvement in computer use skills compared to previous Sonnet models.

Key Specifications

Parameters
-
Context
200.0K
Release Date
February 17, 2026
Average Score
66.9%

Timeline

Key dates in the model's history
Announcement
February 17, 2026
Last Update
February 20, 2026
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$3.00
Output (per 1M tokens)
$15.00
Max Input Tokens
200.0K
Max Output Tokens
64.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-bench Verified
SWE-bench Verified — benchmark for evaluation abilities model solve real tasks from GitHub- • Self-reported
79.6%

Reasoning

Logical reasoning and analysis
GPQA
GPQA Diamond — benchmark for evaluation abilities model answer on questions level PhD by and • Self-reported
89.9%

Other Tests

Specialized benchmarks
ARC-AGI v2
ARC-AGI v2 — benchmark for evaluation abilities to reasoning and generalization • Self-reported
58.3%
MMMLU
MMMLU — version MMLU for evaluation knowledge model on languages • Self-reported
89.3%
CharXiv-R
CharXiv-R — benchmark for evaluation abilities model understand and reason about and • Self-reported
74.7%
MMMU-Pro
MMMU-Pro — version MMMU for evaluation on level experts • Self-reported
75.6%
HLE
HLE (Humanity's Last Exam) — benchmark from questions, experts for verification knowledge AI • Self-reported
49.0%
SimpleQA
SimpleQA — benchmark for evaluation actual accuracy answers model on simple questions. • Self-reported
72.5%
BrowseComp
Agentic search (BrowseComp). • Self-reported
74.7%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, high effort; Pass@1 30% ± 4%. • Verified
30.0%
Finance Agent
Agentic financial analysis (Finance Agent v1.1). • Self-reported
63.3%
Finance Agent v2
• Verified
51.0%
Legal Agent Benchmark
All-pass rate (Harvey held-out set) • Verified
5.4%
LiveBench
2026-01-08, Thinking Medium Effort • Verified
75.5%
MCP Atlas
Scaled tool use (MCP-Atlas). • Self-reported
61.3%
OSWorld
Agentic computer use (OSWorld-Verified). • Self-reported
72.5%
Tau2 Retail
Agentic tool use evaluation (τ2-bench Retail). • Self-reported
91.7%
Tau2 Telecom
Agentic tool use evaluation (τ2-bench Telecom). • Self-reported
97.9%
Terminal-Bench 2.0
Agentic terminal coding (Terminal-Bench 2.0). • Self-reported
59.1%

License & Metadata

License
proprietary
Announcement Date
February 17, 2026
Last Updated
February 20, 2026

Articles about Claude Sonnet 4.6

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.