All News
anthropicclaudebenchmarkssimplebench

Three models now sit above SimpleBench's human baseline, and OpenAI's flagship is not one of them

Claude Opus 5.5 scored 88.4% on SimpleBench, taking rank 1 from Fable 5.1. Three models now clear the 83.7% human baseline — OpenAI's flagship does not.

Vlad MakarovVlad Makarovreviewed and published
3 min read
Three models now sit above SimpleBench's human baseline, and OpenAI's flagship is not one of them

Claude Opus 5.5 has taken the top row of SimpleBench. The Anthropic flagship, launched September 22, scored 88.4% on the benchmark's public leaderboard, displacing Claude Fable 5.1 (86.6%) — the model that in early September became the first of any lab to clear the benchmark's non-specialized human baseline. That write-up called the crossing a single step for the field. Two and a half weeks later it is a queue.

Three above the line, not one

The top of the table now reads:

  • Claude Opus 5.5 (Anthropic): 88.4% — rank 1, new entry
  • Claude Fable 5.1 (Anthropic): 86.6% — rank 2
  • GPT-6 Astra Pro (OpenAI): 86.5% — rank 3
  • Human Baseline*: 83.7% — mean of nine participants

That is the headline number: three models clear 83.7% where on September 5 exactly one did. The sharper detail sits one row lower. GPT-6 Astra, OpenAI's flagship, posts 83.6% — a tenth of a point under the baseline, and under the Pro variant of its own family, which clears it by 2.8 points. Anthropic now holds ranks 1 and 2. Further down, the page added GPT-6 Sol at 73.1% (15th), DeepSeek V4.1 Flash at 66.7% (21st) and Qwen 3.8 27B at 60.2% (36th).

What the page does not say

The site's own introduction has not moved with its table. It still reads: "a non-specialized human baseline is 83.7%, based on our small sample of nine participants, outperforming every tested LLM, including today's top model, Claude Fable, which scored 81.9%." The title above that paragraph is unchanged as well — "Where Everyday Human Reasoning Still Surpasses Frontier Models". The leaderboard was refreshed; the copy around it survived a second crossing untouched.

One footnote under the table is the only methodology offered: temperature 0.7 and top-p 0.95, except the o1 series, with every score an AVG@5. It does not say who ran the new rows, so treat their provenance as unstated rather than vendor-run or independent. Anthropic's own Opus 5.5 launch page never mentions SimpleBench; its results table runs from Terminal-Bench 4.0 to Chartography. The human figure is the mean of nine test takers — the team's word is "baseline", not "average" — and its strongest member scored 95.4%, still 7 points clear of Opus 5.5.

What's next

An r/singularity thread on the score, posted September 24, drew roughly 730 upvotes and more than 130 comments, and the headline claim there again outran the evidence. Opus 5.5's margin over Fable 5.1 is 1.8 points, on a benchmark whose authors describe trick questions as the point. Whether a fourth model clears the line, whether OpenAI's flagship joins its Pro sibling, and whether any of it replicates outside the SimpleBench team's own harness are the open questions.

Related Articles

Scroll down

to load the next article