All News
anthropicclaude-fable-5-1simplebenchbenchmarks

Claude Fable 5.1 is the first model above SimpleBench's 83.7% human baseline

Claude Fable 5.1 scored 86.6% on SimpleBench, clearing the benchmark's 83.7% non-specialized human baseline — the first model of any lab to cross that line.

Vlad MakarovVlad Makarovreviewed and published
2 min read
Mentioned models
Claude Fable 5.1 is the first model above SimpleBench's 83.7% human baseline

Claude Fable 5.1 has become the first model of any lab to sit above SimpleBench's human baseline. The Anthropic flagship, released September 1, now leads the benchmark's public leaderboard at 86.6% — past the 83.7% mark set by a nine-person, non-specialized human sample. The crossing drew a roughly 600-upvote r/singularity thread on September 3.

Why this benchmark is the line

SimpleBench exists to catch models overcomplicating simple questions: more than 200 multiple-choice items on spatial and temporal reasoning, social intelligence, and what its authors call trick questions — items that read like small stories, with mundane answers models keep reasoning themselves past. When it appeared in late 2024, the technical report showed OpenAI's then-flagship o1-preview at 41.7% against the same 83.7% human baseline. Progress came slowly — Claude Fable was the first model past 80% in June at 81.9%. Fable 5.1, covered at launch, cleared the baseline within days of shipping, with Gemini 3.8 Flash (82.4%) and Muse Spark 1.3 (81.8%) close behind:

  • Claude Fable 5.1 (Anthropic): 86.6% — rank 1, new entry
  • Human baseline: 83.7% — nine participants
  • Gemini 3.8 Flash (Google): 82.4% — rank 2, new entry
  • Claude Fable (Anthropic): 81.9% — rank 3

What the crossing does — and does not — mean

Read it as one benchmark cleared, not human-level reasoning achieved. AVG@5 is the average of five model runs at a fixed temperature, and the human number is the mean of nine test takers — a small, self-selected sample whose strongest member scored 95.4%. The thread's headline said "human average"; the benchmark team's label is "baseline" — with nine test takers, the difference is worth keeping. The result also arrived without Anthropic's own announcement: its launch table, the one r/singularity met with "what are these benchmarks", runs from Terminal-Bench-Science to CursorBench and never mentions SimpleBench. Even the benchmark's homepage still carries its founding claim that every tested LLM trails non-specialized humans — a sentence written while Claude Fable's 81.9% led the table. The update outran the site's own copy.

What's next

The line itself will move. SimpleBench's authors say they are functionalizing questions so memorization helps less, and OpenAI's GPT-6 Astra — launched this week — has not posted a row yet. If it clears the baseline too, the benchmark's role flips from humbling frontier models to measuring the gap to its strongest humans. Whether the rest of the field follows Fable 5.1 across the baseline, and how much of the gap to the 95.4% top human score closes first, is the open question. Independent re-runs and a larger human sample would settle how much of this crossing is real progress.

Related Articles

Scroll down

to load the next article