GPT-6.1 Sol tops Epoch's FrontierMath Tier 4 at 100%, on a board that prints no error bar for it
Epoch AI's FrontierMath Tier 4 lists GPT-6.1 Sol at a perfect 100.0%, above GPT-6 Astra's 97.6%. We read the leaderboard, and the interval it does not print.

On October 2, a thread on r/singularity pointed at a new number on Epoch AI's FrontierMath Tier 4 (v2) leaderboard: GPT-6.1 Sol, run at maximum reasoning effort, now sits at 100.0%, the only perfect row on a board of 70 tested models. The previous best was GPT-6 Astra's 97.6%. Epoch is an independent evaluator rather than the model's vendor, and it ran these tests itself — the hub separates "Epoch AI internal runs" from external submissions. That provenance is the reason the figure is worth reading at all.
OpenAI shipped GPT-6.1 Sol on September 29 as a cheaper sibling of GPT-6 Astra, marketed as near-Astra intelligence for a fifth of the price. Astra had already claimed 97.6% on Tier 4 in its own materials, and that is precisely the value Epoch now shows for Astra — a vendor self-report and a third-party run meeting on the same figure.
The board, and the intervals it prints
Epoch released this version of the benchmark on June 12, 2026. The table below takes the top of the leaderboard as Epoch published it, with the accuracy figure and the confidence interval printed beside each row.
| Model (setting) | Accuracy | Interval |
|---|---|---|
| GPT-6.1 Sol (max) | 100.0% | none printed |
| GPT-6 Astra (high / xhigh / max) | 97.6% | ±2.4 |
| Claude Opus 5.5 (max) | 95.0% | ±4.1 |
| Claude Fable 5 (max) | 90.2% | ±4.6 |
| GPT-6 Sol (max) | 90.0% | ±5.2 |
| GPT-6 Astra (low) | 87.8% | ±5.2 |
| GPT-6 Astra (none) | 82.9% | ±5.9 |
| Google DeepMind AI co-mathematician | 75.6% | ±6.7 |
Read the two columns together and the ordering loses some of its teeth. Every model below the top row carries an interval of 2.4 to 7.7 points. Astra's 97.6% comes with a ±2.4, which reaches 100.0% at the top of its range — the same 2.4 points that separate it from first place.
Those bands are the leaderboard's most useful feature. A model run over a fixed problem set and a fixed attempt budget produces a point estimate whose spread comes from sampling, scaffolding and scoring choices, so Epoch publishes the spread rather than a bare number. The bands let a reader see when a ranking is a real gap and when it is noise. Here, the top of the board falls into the second category.
What a perfect score does not carry
The 100.0% row has no interval. Benchmarks usually attach one because a finite set of problems, a fixed number of attempts and a scored rubric all introduce variance, and Epoch prints such a band beside every other row it lists. A perfect score has none to print, or none published, and the leaderboard does not say which. The gap the row implies is real; the uncertainty around it is invisible.
That asymmetry matters more than the two-point headline gap. If Astra's interval reaches 100.0% at its top end and its point estimate is 97.6%, the board cannot cleanly establish that GPT-6.1 Sol is a better mathematician than Astra — only that its mean of sampled runs came in higher. On a set this small, a top-two ordering that sits inside a shared error band is a coin that landed heads, not a settled capability gap.
One more number on Epoch's model page for GPT-6.1 Sol complicates the story rather than confirming it. The same harness scores the model at 94% on FrontierMath Tiers 1-3 (v2), the easier base set, and 100% on Tier 4. The model is not perfect everywhere; it is perfect on the 43 hardest problems. The leaderboard shows the result without explaining the shape, and the version number in the benchmark's name is a reminder that Epoch revised it in June, addressing errors in 42% of its problems.
Epoch's other numbers on the same page keep the profile mixed. GPT-6.1 Sol reaches 94% on FrontierMath Tiers 1-3, 100% on OTIS Mock AIME 2024-2025 and 95% on GPQA Diamond. But it trails Claude Opus 5.5 on Text Arena Coding, 1759 against 1820, and on Furniture Assembly, 80% against 83%; its Chess Puzzles score of 61% sits well below that benchmark's best of 72%. One perfect row on the hardest math tier does not make the model a top scorer everywhere, which is the argument for reading a leaderboard rather than a headline.
Two of the forty-three problems are public
The full dataset is 338 problems: a base set of 295 across Tiers 1-3, plus an expansion of 43 exceptionally difficult problems called Tier 4. Of the twelve problems Epoch makes public, ten come from Tiers 1-3 and two from Tier 4. That means 41 of the 43 problems behind the perfect score cannot be reproduced outside Epoch's administered harness — the scoring code, the attempt budget and the rubric all live with Epoch.
This is a design choice, not a scandal. Holding the problems back is what keeps them from leaking into training data, which is the point of a frontier eval, and it is the same trade-off Epoch's open-problems board manages through labeled results — its board now marks some solves as "human + AI" rather than autonomous, as covered in our earlier piece on that shift. But it also means a 43-problem tier is small enough that a perfect run is far less surprising than the same result on 338, and that an outsider can check the headline only by trusting the evaluator's transcript.
What would settle it
Three things would move this from a published number to a checked one. A third party running the Tier 4 problems outside Epoch's harness would test whether 100% survives a different scaffold. Epoch publishing more of the 43 problems, even a handful more, would let more than two of them be attempted independently. And a second vendor's model cleared through the same tier would show whether saturation at 100% is a property of the benchmark or of one system.
None of that is in the leaderboard, and none of it is required for the figure to be real. Epoch measured what it measured and reported a 100% at the top of a benchmark it controls, a genuinely new result on a set that had topped out at 97.6%. The open questions are about the width of that result, not its existence.


