Claude Opus 4.8
MultimodalClaude Opus 4.8 is Anthropic's upgrade to Opus 4.7 and its most capable general-access model at release, with improvements across software engineering, agentic tool use, reasoning, computer use, and knowledge-work benchmarks while shipping at the same price ($5/$25 per million input/output tokens).
Key Specifications
Parameters
-
Context
1.0M
Release Date
May 28, 2026
Average Score
72.8%
Timeline
Key dates in the model's history
Announcement
May 28, 2026
Last Update
August 27, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$5.00
Output (per 1M tokens)
$25.00
Max Input Tokens
1.0M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
Standard harness with thinking blocks included in sampling, averaged over 5 trials. • Self-reported
Reasoning
Logical reasoning and analysis
GPQA
GPQA Diamond. Averaged over 25 trials. • Self-reported
Other Tests
Specialized benchmarks
BrowseComp
Single-agent with web search, web fetch, code execution, adaptive thinking at max effort, 10M-token limit, context compaction at 200k tokens. Multi-agent (orchestrator + blocking subagents) reaches 88.5%. • Self-reported
CharXiv-R
CharXiv Reasoning. Adaptive thinking at max effort with Python tools: 89.9%. Without tools: 80.5%. 1,000 validation questions averaged over 5 runs. • Self-reported
CyberGym
Targeted vulnerability reproduction (pass@1) over 1,507 tasks without deployed safeguards. With Tier-3 safeguards: 1.0%. • Self-reported
DeepSearchQA
F1 score. Web search, web fetch, code execution, max reasoning effort, adaptive thinking, 1M token budget. • Self-reported
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 59% ± 2%. • Self-reported
Finance Agent
Agentic financial analysis (Finance Agent v2). Evaluated by Vals AI with adaptive thinking at max effort. • Self-reported
Finance Agent v2
• Self-reported
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at max effort: 46.5%. • Self-reported
FrontierSWE
Claude Code • Self-reported
Graphwalks BFS >128k
F1 score on the 1M-token subset, averaged over 5 trials. 256K subset: 85.9%. • Self-reported
Graphwalks parents >128k
F1 score on the 1M-token subset, averaged over 5 trials. 256K subset: 99.3%. • Self-reported
HealthBench Professional
Length-adjusted score with adaptive thinking at max effort. Claude Sonnet 4.6 as grader, averaged over 5 trials. • Self-reported
Humanity's Last Exam
Multidisciplinary reasoning at max effort. With tools (web search, web fetch, code execution): 57.9%. Without tools: 49.8%. • Self-reported
Include
Multilingual regional-knowledge benchmark covering 44 languages. Adaptive thinking enabled with structured JSON output. • Self-reported
LiveBench
2026-01-08, Thinking xHigh Effort • Self-reported
MCP Atlas
Real-world MCP tool use across production-like servers. Evaluated by Scale AI (April 2026 config: 100 tool-call budget). Mean claim coverage of 86.2%. • Self-reported
OfficeQA Pro
Internal agentic harness with extracted-text documents and code-execution tools. OfficeQA full set: 77.6%. • Self-reported
OSWorld-Verified
Agentic computer use evaluation. 361 tasks at 100 steps, 1080p resolution, adaptive thinking at max effort. Pass@1 averaged over 5 seeds. Anthropic updated harness methodology (zoom-tool fix, 128K max tokens per turn). • Self-reported
ScreenSpot Pro
GUI grounding. Adaptive thinking at max effort with Python tools: 87.9%. Without tools: 82.3%. Averaged over 5 runs. • Self-reported
SWE-bench Multilingual
300 problems across 9 programming languages. • Self-reported
SWE-Bench Multimodal
Visual context (screenshots, design mockups) added to issue descriptions. Internal harness. • Self-reported
SWE-Bench Pro
Agentic coding evaluation. Standard harness, adaptive thinking at max effort, averaged over 5 trials. • Self-reported
Terminal-Bench 2.0
Terminal-Bench 2.1 with the Terminus-2 public harness on Daytona, high effort, mean reward over 5 attempts per task (89 tasks, 445 trials). • Self-reported
Toolathlon
Pass@1 averaged over 3 trials across all 108 tasks. Internal harness with adaptive thinking at max effort. Pass@3: 67.6%. • Self-reported
License & Metadata
License
proprietary
Announcement Date
May 28, 2026
Last Updated
August 27, 2026
Similar Models
All ModelsClaude Opus 4.6
Anthropic
MM
Best score:1.0 (TAU)
Released:Feb 2026
Price:$5.00/1M tokens
Claude Sonnet 4.6
Anthropic
MM
Best score:0.9 (GPQA)
Released:Feb 2026
Price:$3.00/1M tokens
Claude Sonnet 4.5
Anthropic
MM
Best score:0.9 (TAU)
Released:Sep 2025
Price:$3.00/1M tokens
Claude Opus 4.7
Anthropic
MM
Best score:0.9 (GPQA)
Released:Apr 2026
Price:$5.00/1M tokens
Claude Opus 4.5
Anthropic
MM
Best score:0.9 (TAU)
Released:Nov 2025
Price:$5.00/1M tokens
Claude Sonnet 5
Anthropic
MM
Released:Jun 2026
Price:$2.00/1M tokens
Claude Opus 5
Anthropic
MM
Released:Jul 2026
Price:$5.00/1M tokens
Claude Fable 5
Anthropic
MM
Released:Jun 2026
Price:$10.00/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.