All News
anthropicclaudereleaseagentic-aipricingsafety

Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cache cut and a 20% per-task markup

Anthropic released Claude Fable 5.1 and Mythos 5.1 — one model behind two safeguard tiers. Cache reads fell 75%, yet measured per-task costs rose about a fifth.

Vlad MakarovVlad Makarovreviewed and published
6 min read
Mentioned models
Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cache cut and a 20% per-task markup

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, billing the pair as its most advanced models for coding and knowledge work. Fable 5.1 is generally available across the Claude API, Claude Code, and Claude.ai, plus AWS, Google Cloud, and Azure; Mythos 5.1 reaches only vetted cyber and life-science groups.

The release answers the feedback that followed Fable 5's June debut: price, data retention, and safeguards that blocked legitimate work. Fable 5 already had an adoption problem — its spend share flat near 11% since July, price the usual suspect. Fable 5.1 targets all three complaints.

One model, two doors

Anthropic's blog is explicit about it: "Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards." Community analysts (eliebakouch, nrehiew_) read the pair the same way: identical weights, different safety and routing.

Mythos 5.1's safeguards are tuned for defensive cyber work and professional life-science R&D. On Terminal-Bench 4.0 it scores 60.9% to Fable 5.1's 55.8%; Anthropic blames its older cyber safeguards for the gap and expects it to shrink. Access runs through the Cyber Verification Program and a Life Sciences program built with the US government, whose first participants are enrolled. For now Mythos 5.1 reaches only US organizations, with wider access coordinated through the government.

A 75% discount with a catch

List prices did not move — $10/$50 per million input/output tokens, same as Fable 5 — and cache writes stay at $12.50. Cache reads, usually the bulk of a long agent session's bill, drop from $1.00 to $0.25. Anthropic estimates that cuts total cost roughly 25% for typical workloads and up to ~45% for highly agentic, context-heavy work, based on four weeks of August usage at default effort.

Third-party measurement complicates it. Artificial Analysis found Fable 5.1 burns ~1.7x the output tokens of Fable 5 at max effort — 140 million output tokens on its Intelligence Index — so even with cheaper cache, measured cost per task ran ~20% higher: $3.76 at max effort versus $2.34 for Opus 5. Latent Space's recap led with the same split: "75% cache price cut but 70% more output tokens." Anthropic's estimate assumes default effort — High in Claude Code, Medium in Cowork and Claude.ai — where Low and Medium settings reach Fable-5-level results for less. At max effort the model simply talks more. Elsewhere: 1M context, 128K max output, text+image input, adaptive thinking always on.

Benchmarks

All figures below are Anthropic's own, from the launch post and system card: unaudited vendor numbers, production safeguards enabled.

BenchmarkFable 5.1Fable 5Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%24.7%29.0%22.4%
Terminal-Bench 4.055.8%42.0%52.3%37.3%
GDPval-AA v2 (Elo)1853172318241711
OSWorld 2.0 (partial)77.9%72.9%75.4%
Humanity's Last Exam (no tools)60.9%57.8%56.6%
CursorBench 3.2.073.4%70.5%70.0%67.2%

Two eval-report footnotes matter. Terminal-Bench-Science carries a standard error of ±3.5–4.5 points, and the public leaderboard's Opus 5 and Fable 5 numbers sit within noise of Anthropic's re-runs. Safeguards were part of the measurement: where they intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0 (strict: 41.7% for 5.1), and Fable 5 zeroed on AutomationBench, where 5.1 posted 31.4% to 17.1%. Cyber tasks were done by Opus 4.8, biology by Opus 5 — lowering the scores, Anthropic says.

Independent numbers published within hours largely agree, with caveats. Artificial Analysis scores Fable 5.1 at 66 on its Intelligence Index at max effort — ahead of Opus 5 (63), Fable 5 (62), and GPT-5.6 Sol (61) — with 59.1% on HLE, 91.4% on Terminal-Bench v2.1, and 62.0% on SciCode. AA cautions it is effectively tied with Opus 5 on several agentic knowledge-work measures, and its run used Anthropic's default fallback, routing ~4% of output tokens to Opus 4.8 or Opus 5. AA also finds it slow and verbose: 66.2 output tokens per second versus a 70-token median.

The safeguards reset

Enterprise Frontier Safeguards (EFS) is the structural change: customers keep data on infrastructure they control, with human review defaulting to the customer — zero-data-retention privacy without giving up misuse monitoring. Anthropic built EFS with 100+ customers and AWS, Google Cloud, and Azure; it rolls out in phases from this fall across Claude Code, Claude Enterprise, Bedrock, and Foundry. Until then, eligible customers run Fable 5.1 and Fable 5 with zero data retention.

The safeguards themselves got more surgical. Cyber interventions in Claude Code drop ~60% per session: Fable 5.1 may now find software vulnerabilities but not develop exploits, and dual-use tasks — penetration testing, exploit generation, binary-based scanning — still route to Opus. Biology safeguards fire 85% less often on benign medical and elementary questions, while life-science R&D queries keep going to Opus; the Mythos program is the sanctioned route past that gate. Anthropic also closed a documented distillation technique: new API accounts can no longer edit prior context while preserving the model's thinking transcript — a change that extends to all users with future releases. Its own risk review finds Mythos 5.1's cyber capabilities the strongest Anthropic has released, still within the lower tier of its Frontier Compliance Framework, with biology above Mythos 5 but below the next Responsible Scaling Policy tier.

What early partners say

Cognition's Walden Yan: "We're moving our Opus 5 traffic in Devin to Claude Fable 5.1 on launch day... a Fable-class model is finally economical for the workloads we'd kept on Opus." Millennium's Damien describes Fable 5.1 finally explaining a crash "about one in a million runs, that nobody on our team had explained in four to five years" — it disassembled an external vendor library and matched it against the core dump. Every's Dan Shipper calls it "friendly Fable... about twice as fast as Opus 5 and used half as many tokens." MongoDB's Ron Sanzone: the prototype "ran for hours unattended... I would wake up in the morning to the next phase finished." Jane Street reports state-of-the-art results on trading intuition.

The bottom line

Fable 5.1 is Anthropic's answer to the charge that its flagship was a supergenius in a datacenter enterprises could not practically use: cheaper re-reads for agentic workloads, a real privacy architecture, fewer safeguard interruptions. Whether it is cheaper per finished task is unresolved and regime-dependent: default effort with heavy cache reuse points down; max effort cost ~20% more, where AA calls Opus 5 the value pick at $2.34 per task. The vendor numbers remain unaudited at these settings. What would settle it: third-party per-task telemetry at production effort levels, not another list-price comparison.

Related Articles

Scroll down

to load the next article