Anthropic logo

Claude Opus 5.5

Multimodal
Anthropic

Claude Opus 5.5 is the first model in Anthropic's Claude 5.5 family. It performs at Claude Fable 5.1 level on most work at about 40% lower cost than Claude Opus 5, and targets long-running agentic coding and knowledge work. It accepts text and image input with text output, supports multilingual and vision workloads, tool use, and adaptive thinking that is always on (API default effort medium). Context window is 1M tokens with 128K max output on the sync Messages API; reliable knowledge and training-data cutoffs are June 2026. First-party pricing is $4/$20 per million input/output tokens, with cache reads at $0.20 and 5-minute cache writes at $5 per million tokens. Fast mode is available in Claude Code and on the Claude Platform at $8/$40 per million input/output tokens with up to 2.5x speed (same model id; not a separate deployment). Capability scores are self-reported from the Claude Opus 5.5 System Card (Table 8.1.A and §8) and the launch announcement. Available on the Claude API as `claude-opus-5-5`.

Key Specifications

Parameters
-
Context
1.0M
Release Date
September 22, 2026
Average Score
71.1%

Timeline

Key dates in the model's history
Announcement
September 22, 2026
Last Update
September 23, 2026
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
June 1, 2026
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$4.00
Output (per 1M tokens)
$20.00
Max Input Tokens
1.0M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
AA-Briefcase v1.1
System Card Table 8.1.A. AA-Briefcase v1.1 Elo at max effort (max_score 3000). • Self-reported
60.7%
ArXivMath
System Card §8.9. ArXivMath August 2026 (57 problems); without tools; max effort; averaged over four attempts per problem. Evaluated internally without safeguards classifiers. • Self-reported
91.2%
AutomationBench
Opus 5.5 with production safeguards enabled. Zapier AutomationBench; runs without fallback models, so safeguard interventions counted as failures. Adaptive thinking at max effort. • Self-reported
40.0%
BenchCAD
System Card §8.13.2. BenchCAD Vision2Code 1k subset; voxel IoU without tools; adaptive thinking at max effort; averaged over five runs. • Self-reported
73.0%
BenchCAD (with Python tool)
System Card §8.13.2. BenchCAD Vision2Code 1k subset; voxel IoU with tools; adaptive thinking at max effort; averaged over five runs. • Self-reported
96.2%
BioMysteryBench
System Card §8.17.1. BioMysteryBench Human Solvable subset 89.3%; adaptive thinking at max effort; bash/file tools with package allow-list. • Self-reported
89.3%
Chartography
Opus 5.5 with production safeguards enabled. Chartography with tools; adaptive thinking at max effort. • Self-reported
89.0%
CursorBench 4.0
Opus 5.5 with production safeguards enabled. CursorBench 4.0; adaptive thinking at max effort. • Self-reported
57.8%
DeepSWE 1.1
System Card §8.3. DeepSWE v1.1; 5-trial average. Adaptive thinking at max effort. • Self-reported
74.2%
FrontierCode 1.1
Opus 5.5 with production safeguards enabled. FrontierCode v1.1 Main subset; adaptive thinking at max effort. • Self-reported
54.4%
FrontierSWE V2
System Card §8.7. FrontierSWE v2; Proximal harness; max effort; 5 trials per task; mean across trials. • Self-reported
62.3%
GDPval-AA 2.1
Opus 5.5 with production safeguards enabled. GDPval-AA v2.1 Elo (max_score 3000); adaptive thinking at max effort. • Self-reported
61.5%
Global-MMLU
System Card §8.16.1. GMMLU average accuracy 94.3% across 42 languages; adaptive thinking at max effort; single trial; no tools. • Self-reported
94.3%
HealthBench
System Card §8.15.1. HealthBench length-adjusted score 60.6% (penalizes verbose responses); adaptive thinking at max effort; 5-trial average. Prefer length-adjusted over raw 68.1% to match other catalog models and avoid duplicate benchmark_id records. • Self-reported
60.6%
HealthBench Professional
System Card Table 8.1.A length-adjusted score 65.6% (prefer Table 8.1.A over §8.15.2 raw 77.1%). Adaptive thinking at max effort; 5-trial average; safety classifiers enabled with refusal fallback to Claude Opus 5. Length-adjusted per HealthBench Professional paper method. • Self-reported
65.6%
Humanity's Last Exam (no tools, text-only)
System Card Table 8.1.A / §8.11.1. HLE no tools (reasoning-only); adaptive thinking at max effort. Alongside with-tools 67.7% already cataloged from the announcement table. • Self-reported
64.4%
Humanity's Last Exam (with tools, text-only)
Opus 5.5 with production safeguards enabled. Humanity's Last Exam with tools; adaptive thinking at max effort. • Self-reported
67.7%
LatchBio SingleCellBench
System Card §8.17.2. LatchBio SingleCellBench 61.2%; adaptive thinking at max effort; bash/file tools. • Self-reported
61.2%
LatchBio SpatialBench Verified
System Card §8.17.2. LatchBio SpatialBench Verified 72.0%; adaptive thinking at max effort; bash/file tools. • Self-reported
72.0%
Legal Agent Benchmark
System Card §8.14.2. Legal Agent Benchmark held-out set; all-pass rate 8.3% at max effort (AA harness). Mean criterion-pass rate 91.2% noted in card but not stored as a separate catalog metric (all-pass is the standard LAB reporting id). • Self-reported
8.3%
MILU
System Card §8.16.2. MILU average accuracy 93.1% across 11 languages; adaptive thinking at max effort; averaged over 5 trials; no tools. • Self-reported
93.1%
OfficeQA
System Card §8.14.1. OfficeQA full set; max effort; mean of five runs; agentic harness with extracted-text documents and code-execution tools. • Self-reported
78.9%
OfficeQA Pro
System Card §8.14.1. OfficeQA Pro (133-question subset); max effort; mean of five runs. • Self-reported
67.7%
OSWorld 2.0
Opus 5.5 with production safeguards enabled. OSWorld 2.0 partial score (announcement reports partial, not strict); adaptive thinking at max effort. • Self-reported
81.8%
Program Bench
System Card §8.10.1. ProgramBench 166 golden tasks; mini-swe-agent harness; no 6-hour time limit; hidden test pass rate against tests the reference binary passes. • Self-reported
91.2%
SWE-bench Multilingual
System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. 300 problems across nine programming languages. • Self-reported
93.9%
SWE-Bench Multimodal
System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. Visual context (screenshots and design mockups). • Self-reported
61.4%
SWE-Bench Pro
System Card Table 8.1.A / §8.2. Adaptive thinking at max effort; 5-trial average. Production safeguards as in card capability summary. • Self-reported
89.9%
Terminal-Bench 4.0
Opus 5.5 with production safeguards enabled. Terminal-Bench 4.0 accuracy 66.4% at xhigh effort (Claude Code harness). Unless otherwise noted, other Opus 5.5 results on this announcement table use adaptive thinking at max effort; TB4 is explicitly reported at xhigh. • Self-reported
66.4%
Terminal-Bench-Science 0.1
Opus 5.5 with production safeguards enabled. Terminal-Bench-Science 0.1; adaptive thinking at max effort. • Self-reported
58.7%
Toolathlon-Verified
System Card §8.14.5. Toolathlon Verified Pass@1 77.8% (3-trial average across 108 tasks). Adaptive thinking at max effort; safety/sandbox stops counted as failures. • Self-reported
77.8%

License & Metadata

License
proprietary
Announcement Date
September 22, 2026
Last Updated
September 23, 2026

Compare Claude Opus 5.5

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.