Anthropic logo

Claude Fable 5

Multimodal
Anthropic

Claude Fable 5 is Anthropic's generally available deployment of the same underlying weights as the restricted Claude Mythos 5 model, adding production safeguards for public use. It leads on agentic coding, software engineering, and computer-use benchmarks such as SWE-Bench Verified, FrontierSWE, and OSWorld-Verified while offering a 1M-token context window.

Key Specifications

Parameters
-
Context
1.0M
Release Date
June 9, 2026
Average Score
60.7%

Timeline

Key dates in the model's history
Announcement
June 9, 2026
Last Update
August 27, 2026

Technical Specifications

Parameters
-
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$10.00
Output (per 1M tokens)
$50.00
Max Input Tokens
1.0M
Max Output Tokens
128.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Programming

Programming skills tests
SWE-Bench Verified
Fable 5 production-safeguarded configuration on the 500-problem SWE-Bench Verified subset; average over 5 trials with thinking blocks included. Mythos 5 scored 95.5%.Self-reported
95.0%

Other Tests

Specialized benchmarks
AutomationBench
Zapier AutomationBench private held-out set at max effort. End-to-end business workflows across simulated company APIs.Self-reported
17.4%
BioMysteryBench
Hard subset: 46.1%. Human-solved subset: 83.9%. Blocking safeguards cause Fable 5 to perform closer to Opus 4.8 due to fallbacks. System card reports this as a Mythos 5 or combined launch-table score rather than a separate Fable-only production score; Fable may fall back to Opus 4.8 when safeguards trigger.Self-reported
46.1%
Blueprint-Bench 2
Spatial reasoning evaluation.Self-reported
38.6%
DeepSWE 1.1
DeepSWE v1.1 leaderboard, mini-swe-agent harness, max effort; Pass@1 70% ± 4%.Self-reported
70.0%
ExploitBench
Cybersecurity capture rate (Cap%). Blocking safeguards cause Fable 5 to perform closer to Opus 4.8 due to fallbacks; score reflects Claude Mythos 5. System card reports this as a Mythos 5 or combined launch-table score rather than a separate Fable-only production score; Fable may fall back to Opus 4.8 when safeguards trigger.Self-reported
78.0%
Finance Agent v2
Self-reported
56.3%
FrontierCode
FrontierCode Main subset at xhigh reasoning effort. Score 46.3%, pass rate 48.8%; Fable 5 ranks #1 and outperforms every other model even at medium effort.Self-reported
46.3%
FrontierCode 1.1
FrontierCode 1.1 current leaderboard; mergeability score at xhigh effort: 53.5%.Self-reported
53.5%
FrontierSWE
Claude CodeSelf-reported
90.0%
GDP.pdf
Document and vision-heavy knowledge-work evaluation, no tools.Self-reported
29.8%
HealthBench Professional
Health evaluation. Blocking safeguards cause Fable 5 to perform closer to Opus 4.8 due to fallbacks. System card reports this as a Mythos 5 or combined launch-table score rather than a separate Fable-only production score; Fable may fall back to Opus 4.8 when safeguards trigger.Self-reported
66.0%
Humanity's Last Exam
Multidisciplinary reasoning. With tools: 64.5%. Without tools: 59.0%. Blocking safeguards cause Fable 5 to perform closer to Opus 4.8 on some questions due to fallbacks. System card reports this as a Mythos 5 or combined launch-table score rather than a separate Fable-only production score; Fable may fall back to Opus 4.8 when safeguards trigger.Self-reported
64.5%
Legal Agent Benchmark
Legal Agent Benchmark, Harvey's held-out set.Self-reported
13.3%
LiveBench
2026-01-08, Thinking xHigh EffortSelf-reported
78.3%
OSWorld-Verified
Fable 5 production-safeguarded configuration. OSWorld score reflects the updated harness with zoom-tool bug fix and 128K max tokens per turn.Self-reported
85.0%
SWE-Bench Pro
Fable 5 production-safeguarded configuration on SWE-Bench Pro; average over 5 trials with thinking blocks included. Mythos 5 scored 80.3%.Self-reported
80.0%
Terminal-Bench 2.1
Fable 5 production-safeguarded configuration on the mini-SWE-agent harness, high effort, 89 tasks x 5 attempts. 20.9% of trials hit a safety refusal and fell back to Claude Opus 4.8 for the rest of the trajectory; Mythos 5 scored 88.0%.Self-reported
84.3%

License & Metadata

License
proprietary
Announcement Date
June 9, 2026
Last Updated
August 27, 2026

Compare Claude Fable 5

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.