Anthropic logo

Claude Opus 4.6

Multimodal
Anthropic

Claude Opus 4.6 is Anthropic's most intelligent model for building agents and coding. Significantly improved coding skills: more thorough planning, sustained agentic task support, reliable performance in large codebases, improved code review and debugging. Context window: 200K tokens by default, 1M tokens available in beta with premium pricing ($10/$37.50 per million input/output tokens at >200K). Up to 128K output tokens. New API features: adaptive thinking (the model decides when to use extended thinking), effort control (low/medium/high/max), context compression for long-running tasks. Leads on Terminal-Bench 2.0 (agentic coding), Humanity's Last Exam (multidisciplinary reasoning), GDPval-AA (knowledge work in finance, law), BrowseComp (information retrieval), DeepSearchQA (deep agentic search). Supports agent teams in Claude Code, Claude in Excel, and Claude in PowerPoint.

Key Specifications

Parameters
-
Context
1.0M
Release Date
February 4, 2026
Average Score
74.8%

Timeline

Key dates in the model's history
Announcement
February 4, 2026
Last Update
February 6, 2026
Today
October 5, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
May 1, 2025
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$5.00
Output (per 1M tokens)
$25.00
Max Input Tokens
1.0M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
SWE-Bench Verified — solution real tasks from GitHub issues. • Self-reported
80.8%

Reasoning

Logical reasoning and analysis
GPQA
Accuracy GPQA Diamond. • Self-reported
91.3%

Other Tests

Specialized benchmarks
Vending-Bench 2
Final in USD. Simulation vending business for year work. Starting $5,000 • Self-reported
100.0%
GDPval-AA
Elo rating. Independent evaluation Artificial Analysis. Outperforms GPT-5.2 on ~144 Elo and Claude Opus 4.5 on 190 points. • Self-reported
53.5%
AIME 2025
Accuracy Consensus@64 (most often occurring answer among 64 samples). Independent evaluation Artificial Analysis. • Self-reported
100.0%
TAU2 Telecom
tool use (τ2-bench Telecom) • Self-reported
99.0%
Graphwalks Parents >128K
GraphWalks Parents 256K 1M. F1 with context 1M, average from 5 attempts • Self-reported
95.0%
MRCR v2 (8-needle)
OpenAI MRCR v2 256K 8-needles. Mean Match Ratio with 1M, average from 5 attempts • Self-reported
93.0%
Humanity's Last Exam
Accuracy on HLE benchmark. • Self-reported
53.1%
BrowseComp
Accuracy BrowseComp — by for search complex information • Self-reported
84.0%
ARC-AGI v2
ARC-AGI-2 — reasoning through • Self-reported
68.8%
CharXiv-R
CharXiv-R — reasoning about scientific from arXiv • Self-reported
77.4%
CyberGym
Cybersecurity vulnerability reproduction. No thinking, default effort, temperature, and top_p. Given a 'think' tool for interleaved thinking in multi-turn evaluations. • Self-reported
73.8%
DeepSearchQA
F1 score. 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks. Programmatic search tools, context compaction, adaptive thinking, max effort up to 10M total tokens. • Self-reported
91.3%
FigQA
LAB-Bench FigQA. With adaptive thinking, image cropping tool, and max effort. Scores averaged over 5 runs. Without image cropping tool (reasoning only): 58.0%. • Self-reported
78.3%
Finance Agent
Agentic financial analysis evaluation. • Self-reported
60.7%
FrontierSWE
Claude Code • Verified
56.0%
Graphwalks BFS >128k
GraphWalks BFS 256K subset of 1M. F1 score with 64k output tokens, 1M context window, average of 5 trials. Max output tokens: 61.1. Full 1M variant: 41.2 (64k output). • Self-reported
61.5%
Graphwalks parents >128k
GraphWalks Parents 256K subset of 1M. F1 score with max output tokens, 1M context window, average of 5 trials. At 64k output tokens: 95.1. Full 1M variant: 72.0 (max), 71.1 (64k). • Self-reported
95.4%
Legal Agent Benchmark
All-pass rate (Harvey held-out set) • Verified
4.2%
LiveBench
2026-01-08, Thinking High Effort • Verified
76.3%
MCP Atlas
Scaled tool use. At high effort, reached industry-leading score of 62.7%. At max effort, score is higher. • Self-reported
62.7%
MMMLU
Multilingual Q&A. • Self-reported
91.1%
MMMU-Pro
Visual reasoning with tools: 77.3%. Without tools: 73.9%. • Self-reported
77.3%
MRCR v2 (8-needle)
MRCR v2 1M 8-needles variant. Mean Match Ratio with max output tokens, average of 5 trials. Anthropic also reports 78.3 with 64k output tokens and 93.0 for the 256K 8-needle subset. • Self-reported
76.0%
OpenRCA
Software failure diagnosis. 1 point if all generated root-cause elements match ground-truth, 0 points if any mismatch. Run on benchmark author's harness, graded using official methodology. • Self-reported
34.9%
OSWorld
Agentic computer use evaluation. • Self-reported
72.7%
SWE-bench Multilingual
Averaged over 25 trials, each run with adaptive thinking, max effort, default sampling settings (temperature, top_p), and with thinking blocks included in sampling results. • Self-reported
77.8%
Tau2 Retail
Agentic tool use evaluation (τ2-bench Retail). • Self-reported
91.9%
Terminal-Bench 2.0
Terminus-2 harness. 1× guaranteed / 3× ceiling resource allocation, 5–15 samples per task across staggered batches. • Self-reported
65.4%

License & Metadata

License
proprietary
Announcement Date
February 4, 2026
Last Updated
February 6, 2026

Articles about Claude Opus 4.6

Anthropic Surpasses OpenAI in Annual Recurring Revenue

Anthropic has reportedly crossed $25B ARR, overtaking OpenAI for the first time. Ramp data shows 73% of new enterprise customers now choose Claude.

7 min

Claude Code's Deny Rules Silently Bypassed After Source Code Leak

Security firm Adversa AI discovers that Claude Code's deny rules are silently disabled when a command contains 50+ subcommands, letting attackers steal credentials.

2 min
Agents of Chaos: When AI Agents Get OS-Level Access, Everything Breaks

Agents of Chaos: When AI Agents Get OS-Level Access, Everything Breaks

A 30-researcher team from Stanford, Harvard and MIT red-teamed autonomous AI agents built on OpenClaw. The results are alarming: data leaks, identity hijacks, and runaway loops.

8 min

Anthropic Accidentally Ships Claude Code's Entire Source Code

A misconfigured npm package exposes 512,000 lines of Claude Code's TypeScript source, revealing feature flags, system prompts, and a hidden Tamagotchi pet.

3 min
Congress Draws a Line on AI Weapons After Anthropic-Pentagon Standoff

Congress Draws a Line on AI Weapons After Anthropic-Pentagon Standoff

Senate Democrats introduce bills to ban autonomous AI weapons and mass surveillance after the Pentagon blacklisted Anthropic. What the legislation says and why it matters.

8 min
Amodei Calls Brockman's $25M Trump Donation 'Evil'

Amodei Calls Brockman's $25M Trump Donation 'Evil'

In a leaked internal letter, Anthropic CEO Dario Amodei condemned OpenAI president Greg Brockman's $25 million donation to a Trump super PAC and accused rivals of 'dictator-style praise.'

3 min
Anthropic May Have Had an Architectural Breakthrough

Anthropic May Have Had an Architectural Breakthrough

Weeks-old rumors of a fundamental AI architecture discovery now point to Anthropic, fueled by the Mythos leak showing capabilities far beyond incremental scaling improvements.

3 min
Anthropic Accidentally Leaked Its Most Dangerous AI Model

Anthropic Accidentally Leaked Its Most Dangerous AI Model

Security researchers found nearly 3,000 unpublished Anthropic documents in a public data cache — including details on Claude Mythos, a new model with unprecedented cybersecurity capabilities.

4 min
Judge Blocks Pentagon's 'Supply Chain Risk' Label on Anthropic

Judge Blocks Pentagon's 'Supply Chain Risk' Label on Anthropic

A federal judge in San Francisco barred the Department of Defense from designating Anthropic as a supply chain risk, calling the move 'arbitrary and capricious.'

3 min
Every Frontier AI Model Scored Under 1% on ARC-AGI-3. Humans Got 100%.

Every Frontier AI Model Scored Under 1% on ARC-AGI-3. Humans Got 100%.

Chollet's new benchmark drops the same week Jensen Huang declared AGI. GPT-5.4 scored 0.26%. Claude Opus 4.6 scored 0.25%. The gap with humans is 99+ points.

9 min

Chollet to Jensen Huang: 'You Don't Have AGI. You Have Expensive Autocomplete.'

The ARC-AGI creator fires back at NVIDIA's AGI claims with a benchmark that every frontier model fails. His argument is simple: if humans can do it cold, AGI should too.

2 min
Claude Can Now Control Your Computer. Anthropic Says Trust It — Mostly.

Claude Can Now Control Your Computer. Anthropic Says Trust It — Mostly.

Anthropic shipped computer use for Claude Code and Cowork — mouse, keyboard, browser, files. Plus a new Auto Mode that skips permission prompts. macOS only.

7 min

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.