Anthropic logo

Claude Opus 4.8

Multimodal
Anthropic

Claude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens).

Key Specifications

Parameters
-
Context
1.0M
Release Date
May 28, 2026
Average Score
72.8%

Timeline

Key dates in the model's history
Announcement
May 28, 2026
Last Update
August 27, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$5.00
Output (per 1M tokens)
$25.00
Max Input Tokens
1.0M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
Standard harness with thinking blocks included in sampling, averaged over 5 trials.Self-reported
88.6%

Reasoning

Logical reasoning and analysis
GPQA
GPQA Diamond. Averaged over 25 trials.Self-reported
93.6%

Other Tests

Specialized benchmarks
BrowseComp
Single-agent with web search, web fetch, code execution, adaptive thinking at max effort, 10M-token limit, context compaction at 200k tokens. Multi-agent (orchestrator + blocking subagents) reaches 88.5%.Self-reported
84.3%
CharXiv-R
CharXiv Reasoning. Adaptive thinking at max effort with Python tools: 89.9%. Without tools: 80.5%. 1,000 validation questions averaged over 5 runs.Self-reported
89.9%
CyberGym
Targeted vulnerability reproduction (pass@1) over 1,507 tasks without deployed safeguards. With Tier-3 safeguards: 1.0%.Self-reported
78.8%
DeepSearchQA
F1 score. Web search, web fetch, code execution, max reasoning effort, adaptive thinking, 1M token budget.Self-reported
93.1%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 59% ± 2%.Self-reported
59.0%
Finance Agent
Agentic financial analysis (Finance Agent v2). Evaluated by Vals AI with adaptive thinking at max effort.Self-reported
53.9%
Finance Agent v2
Self-reported
53.9%
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at max effort: 46.5%.Self-reported
46.5%
FrontierSWE
Claude CodeSelf-reported
75.0%
Graphwalks BFS >128k
F1 score on the 1M-token subset, averaged over 5 trials. 256K subset: 85.9%.Self-reported
68.1%
Graphwalks parents >128k
F1 score on the 1M-token subset, averaged over 5 trials. 256K subset: 99.3%.Self-reported
83.3%
HealthBench Professional
Length-adjusted score with adaptive thinking at max effort. Claude Sonnet 4.6 as grader, averaged over 5 trials.Self-reported
55.8%
Humanity's Last Exam
Multidisciplinary reasoning at max effort. With tools (web search, web fetch, code execution): 57.9%. Without tools: 49.8%.Self-reported
57.9%
Include
Multilingual regional-knowledge benchmark covering 44 languages. Adaptive thinking enabled with structured JSON output.Self-reported
87.6%
LiveBench
2026-01-08, Thinking xHigh EffortSelf-reported
77.2%
MCP Atlas
Real-world MCP tool use across production-like servers. Evaluated by Scale AI (April 2026 config: 100 tool-call budget). Mean claim coverage of 86.2%.Self-reported
82.2%
OfficeQA Pro
Internal agentic harness with extracted-text documents and code-execution tools. OfficeQA full set: 77.6%.Self-reported
66.2%
OSWorld-Verified
Agentic computer use evaluation. 361 tasks at 100 steps, 1080p resolution, adaptive thinking at max effort. Pass@1 averaged over 5 seeds. Anthropic updated harness methodology (zoom-tool fix, 128K max tokens per turn).Self-reported
83.4%
ScreenSpot Pro
GUI grounding. Adaptive thinking at max effort with Python tools: 87.9%. Without tools: 82.3%. Averaged over 5 runs.Self-reported
87.9%
SWE-bench Multilingual
300 problems across 9 programming languages.Self-reported
84.4%
SWE-Bench Multimodal
Visual context (screenshots, design mockups) added to issue descriptions. Internal harness.Self-reported
38.4%
SWE-Bench Pro
Agentic coding evaluation. Standard harness, adaptive thinking at max effort, averaged over 5 trials.Self-reported
69.2%
Terminal-Bench 2.0
Terminal-Bench 2.1 with the Terminus-2 public harness on Daytona, high effort, mean reward over 5 attempts per task (89 tasks, 445 trials).Self-reported
74.6%
Toolathlon
Pass@1 averaged over 3 trials across all 108 tasks. Internal harness with adaptive thinking at max effort. Pass@3: 67.6%.Self-reported
59.9%

License & Metadata

License
proprietary
Announcement Date
May 28, 2026
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.