Key Specifications
Parameters
-
Context
-
Release Date
April 7, 2026
Average Score
85.1%
Timeline
Key dates in the model's history
Announcement
April 7, 2026
Last Update
August 29, 2026
Technical Specifications
Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$25.00
Output (per 1M tokens)
$125.00
Max Input Tokens
-
Max Output Tokens
-
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
Memorization screens flag a subset of problems. Excluding flagged problems, Mythos Preview's margin over Opus 4.6 holds. • Self-reported
Reasoning
Logical reasoning and analysis
GPQA
GPQA Diamond. • Self-reported
Other Tests
Specialized benchmarks
BrowseComp
Scores higher than Opus 4.6 while using 4.9× fewer tokens. • Self-reported
CharXiv-R
CharXiv Reasoning with tools. Without tools: 86.1%. Opus 4.6: 78.9% (with tools), 61.5% (without tools). • Self-reported
CyBench
100% pass@1. Benchmark saturated. • Self-reported
CyberGym
Cybersecurity vulnerability reproduction. • Self-reported
FigQA
LAB-Bench FigQA with tools. Opus 4.6: 75.1% (with tools). • Self-reported
Graphwalks BFS >128k
GraphWalks BFS 256K–1M. Opus 4.6: 38.7%, GPT-5.4: 21.4%. • Self-reported
Humanity's Last Exam
With tools. Without tools: 56.8%. Anthropic notes Mythos may show some level of memorization at low effort. • Self-reported
MMMLU
Multilingual Q&A. Opus 4.6: 91.1%, Gemini 3.1 Pro: 92.6–93.6%. • Self-reported
OSWorld-Verified
• Self-reported
SWE-bench Multilingual
Memorization screens flag a subset of problems. Excluding flagged problems, Mythos Preview's margin over Opus 4.6 holds. • Self-reported
SWE-Bench Multimodal
Internal implementation. Scores not directly comparable to public leaderboard scores. Opus 4.6: 27.1% on same implementation. • Self-reported
SWE-Bench Pro
Memorization screens flag a subset of problems. Excluding flagged problems, Mythos Preview's margin over Opus 4.6 holds. • Self-reported
Terminal-Bench 2.0
Terminus-2 harness with adaptive thinking at maximum effort, 1M token total task budget. 1× guaranteed / 3× ceiling resource allocation, averaged over 5 attempts per task. 92.1% with 4-hour timeout limits and Terminal-Bench 2.1 updates. • Self-reported
USAMO25
USAMO 2026 math proofs. Opus 4.6: 42.3%, GPT-5.4: 95.2%, Gemini 3.1 Pro: 74.4%. • Self-reported
License & Metadata
License
proprietary
Announcement Date
April 7, 2026
Last Updated
August 29, 2026
Similar Models
All ModelsClaude Opus 4.5
Anthropic
MM
Best score:0.9 (TAU)
Released:Nov 2025
Price:$5.00/1M tokens
Claude Opus 4.6
Anthropic
MM
Best score:1.0 (TAU)
Released:Feb 2026
Price:$5.00/1M tokens
Claude Sonnet 4.6
Anthropic
MM
Best score:0.9 (GPQA)
Released:Feb 2026
Price:$3.00/1M tokens
Claude Opus 4.8
Anthropic
MM
Best score:0.9 (GPQA)
Released:May 2026
Price:$5.00/1M tokens
Claude Opus 4.7
Anthropic
MM
Best score:0.9 (GPQA)
Released:Apr 2026
Price:$5.00/1M tokens
Claude 3.7 Sonnet
Anthropic
MM
Best score:0.8 (GPQA)
Released:Feb 2025
Price:$3.00/1M tokens
Claude 3 Sonnet
Anthropic
MM
Best score:0.9 (ARC)
Released:Feb 2024
Price:$3.00/1M tokens
Claude 3.5 Sonnet
Anthropic
MM
Best score:0.9 (HumanEval)
Released:Oct 2024
Price:$3.00/1M tokens
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.