Anthropic logo

Claude Opus 4.7

Multimodal
Anthropic

Claude Opus 4.7 is Anthropic's latest Opus-class model, a direct upgrade to Opus 4.6 with notable improvements in advanced software engineering, particularly on the most difficult tasks. It handles complex, long-running agentic workflows with rigor and consistency, follows instructions more literally and precisely, and verifies its own outputs before reporting back.

Key Specifications

Parameters
-
Context
1.0M
Release Date
April 16, 2026
Average Score
68.3%

Timeline

Key dates in the model's history
Announcement
April 16, 2026
Last Update
August 27, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$5.00
Output (per 1M tokens)
$25.00
Max Input Tokens
1.0M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
Agentic coding evaluation. Memorization screens flag a subset of problems; excluding those, Opus 4.7's margin over Opus 4.6 holds.Self-reported
87.6%

Reasoning

Logical reasoning and analysis
GPQA
GPQA Diamond. Graduate-level reasoning evaluation.Self-reported
94.2%

Other Tests

Specialized benchmarks
BrowseComp
Agentic search evaluation.Self-reported
79.3%
CharXiv-R
CharXiv Reasoning. Visual reasoning with tools: 91.0%. Without tools: 82.1%.Self-reported
91.0%
CyberGym
Cybersecurity vulnerability reproduction. Opus 4.6's score was updated from the originally reported 66.6 to 73.8 after harness parameter updates to better elicit cyber capability.Self-reported
73.1%
Finance Agent
Agentic financial analysis evaluation (Finance Agent v1.1). State-of-the-art score at release.Self-reported
64.4%
Finance Agent v2
Self-reported
51.5%
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at max effort: 38.5%.Self-reported
38.5%
FrontierSWE
Claude CodeSelf-reported
63.0%
Humanity's Last Exam
Multidisciplinary reasoning. With tools: 54.7%. Without tools: 46.9%.Self-reported
54.7%
Legal Agent Benchmark
All-pass rate (Harvey held-out set)Self-reported
7.1%
LiveBench
2026-01-08, Thinking xHigh EffortSelf-reported
76.9%
MCP Atlas
Scaled tool use evaluation. Opus 4.6's MCP-Atlas score was updated to reflect revised grading methodology from Scale AI.Self-reported
77.3%
MMMLU
Multilingual Q&A.Self-reported
91.5%
OSWorld-Verified
Agentic computer use evaluation.Self-reported
78.0%
SWE-Bench Pro
Agentic coding evaluation.Self-reported
64.3%
Terminal-Bench 2.0
Terminus-2 harness with thinking disabled. 1× guaranteed / 3× ceiling resource allocation averaged over five attempts per task.Self-reported
69.4%

License & Metadata

License
proprietary
Announcement Date
April 16, 2026
Last Updated
August 27, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.