Anthropic logo

Claude Mythos Preview

Multimodal
Anthropic

Claude Mythos Preview is an unreleased general-purpose frontier model from Anthropic, a new tier above Opus (internal codename 'Capybara').

Key Specifications

Parameters
-
Context
-
Release Date
April 7, 2026
Average Score
85.1%

Timeline

Key dates in the model's history
Announcement
April 7, 2026
Last Update
August 29, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$25.00
Output (per 1M tokens)
$125.00
Max Input Tokens
-
Max Output Tokens
-
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
Memorization screens flag a subset of problems. Excluding flagged problems, Mythos Preview's margin over Opus 4.6 holds.Self-reported
93.9%

Reasoning

Logical reasoning and analysis
GPQA
GPQA Diamond.Self-reported
94.6%

Other Tests

Specialized benchmarks
BrowseComp
Scores higher than Opus 4.6 while using 4.9× fewer tokens.Self-reported
86.9%
CharXiv-R
CharXiv Reasoning with tools. Without tools: 86.1%. Opus 4.6: 78.9% (with tools), 61.5% (without tools).Self-reported
93.2%
CyBench
100% pass@1. Benchmark saturated.Self-reported
100.0%
CyberGym
Cybersecurity vulnerability reproduction.Self-reported
83.1%
FigQA
LAB-Bench FigQA with tools. Opus 4.6: 75.1% (with tools).Self-reported
89.0%
Graphwalks BFS >128k
GraphWalks BFS 256K–1M. Opus 4.6: 38.7%, GPT-5.4: 21.4%.Self-reported
80.0%
Humanity's Last Exam
With tools. Without tools: 56.8%. Anthropic notes Mythos may show some level of memorization at low effort.Self-reported
64.7%
MMMLU
Multilingual Q&A. Opus 4.6: 91.1%, Gemini 3.1 Pro: 92.6–93.6%.Self-reported
92.7%
OSWorld-Verified
Self-reported
79.6%
SWE-bench Multilingual
Memorization screens flag a subset of problems. Excluding flagged problems, Mythos Preview's margin over Opus 4.6 holds.Self-reported
87.3%
SWE-Bench Multimodal
Internal implementation. Scores not directly comparable to public leaderboard scores. Opus 4.6: 27.1% on same implementation.Self-reported
59.0%
SWE-Bench Pro
Memorization screens flag a subset of problems. Excluding flagged problems, Mythos Preview's margin over Opus 4.6 holds.Self-reported
77.8%
Terminal-Bench 2.0
Terminus-2 harness with adaptive thinking at maximum effort, 1M token total task budget. 1× guaranteed / 3× ceiling resource allocation, averaged over 5 attempts per task. 92.1% with 4-hour timeout limits and Terminal-Bench 2.1 updates.Self-reported
82.0%
USAMO25
USAMO 2026 math proofs. Opus 4.6: 42.3%, GPT-5.4: 95.2%, Gemini 3.1 Pro: 74.4%.Self-reported
97.6%

License & Metadata

License
proprietary
Announcement Date
April 7, 2026
Last Updated
August 29, 2026

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.