'What are these benchmarks': r/singularity pushes back on Claude Fable 5.1's launch table
r/singularity met Claude Fable 5.1's launch benchmark table with 'what are these benchmarks' — new evals, self-reported scores, and safety routing under fire.

Six minutes after Anthropic's September 1 launch post appeared, a screenshot of Claude Fable 5.1's benchmark table hit r/singularity under the title "What are these benchmarks." The thread — roughly 520 upvotes, more than 170 comments — does not dispute that the model is strong. It disputes the evaluations, and that reception is the story today.
What the community is questioning
The screenshot is Anthropic's own launch table: seven evaluations running from Terminal-Bench-Science 0.1 to CursorBench 3.2.0, with the full numbers covered in our launch piece. The skepticism starts with the evaluation names themselves — Terminal-Bench-Science 0.1, Terminal-Bench 4.0, GDPval-AA v2, AutomationBench — several of which are new or newly versioned, leaving outsiders no independent runs to cross-check against. Commenters compare the table with Fable 5's June launch, ask for third-party replication, and joke about evaluation-name inflation.
Then there are the footnotes, where the harder caveats live:
- Fable 5.1 was evaluated with production safeguards enabled — and where those safeguards intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0 (Fable 5 also zeroed AutomationBench).
- Cyber tasks were completed by Claude Opus 4.8 and biology tasks by Opus 5 — the headline model was not the model doing some of the work.
- GPT-5.6 Sol rows show em-dashes on OSWorld 2.0 and Humanity's Last Exam, with Anthropic citing re-run conditions.
- Standard errors — Terminal-Bench-Science carries ±3.5–4.5 points — appear only in the footnotes.
What third-party testing shows
All of it is vendor self-reported on launch day, so the counterweight is Artificial Analysis, which gave Fable 5.1 an Intelligence Index score of 66 at max effort — the highest it has measured, ahead of Opus 5 (63) and GPT-5.6 Sol (61). Yet even that run carries caveats the community's questions anticipate: the index was evaluated with Anthropic's default fallback, which routed roughly 4% of output tokens to Opus 4.8 or Opus 5 for safety-flagged requests, and Artificial Analysis disclosed supporting pre-release evaluation. Cost is the other live question: at max effort Fable 5.1 costs about 20% more per task than Fable 5, because it emits roughly 1.7x the output tokens; the 75% cache-read cut saves about $1.40 per task but does not fully offset it.
Why the pattern matters
Proliferation of new benchmark versions, launch-day self-reporting, and safety routing have made "state of the art" harder to audit — which is precisely why the demand for real-world, third-party testing keeps rising. What would settle this one: independent re-runs under identical conditions, with logs of which model answered which task, and time for the new evaluations to accumulate runs beyond Anthropic's own.


