All News
epoch-aibenchmarkmathematicsgpt-6-astralean

FrontierMath Erdős: GPT-6 Astra solved 2 of 68 open problems while rivals scored zero

Epoch AI benchmarked models on 68 significant unsolved Erdős problems formalized in Lean. A pre-release GPT-6 Astra solved two. Everyone else scored zero.

Vlad MakarovVlad Makarovreviewed and published
5 min read
FrontierMath Erdős: GPT-6 Astra solved 2 of 68 open problems while rivals scored zero

Epoch AI unveiled FrontierMath Erdős this week, a benchmark built from 68 significant open problems posed by Paul Erdős, the prolific Hungarian mathematician, and its first results produced a headline ready-made for r/singularity: a pre-release version of GPT-6 Astra scored 3%, the only nonzero score among five tested models. The thread, titled "GPT-6 Astra gets 3% on the FrontierMath Erdős Benchmark, while every other Model (that was tested) got 0%," gathered roughly 650 points and more than 100 comments. Both parts of that sentence are true, and both need a footnote: 3% is two problems, and the two solutions are genuine results, verified by machine, while the 0% column hides that every other attempt simply ran out of budget without a proof.

A catalog of 1,217 problems, trimmed to 68

"Erdős problem" sounds like a uniform category, but it covers everything from throwaway exercises to questions that defined whole subfields. Thomas Bloom, the mathematician behind erdosproblems.com, has catalogued Erdős's scattered problems since 2023; Epoch's announcement counts 1,217 of them, with 652 still open as of August 2026. From that unsolved set Bloom selected his favorites, problems he judged both significant and difficult: 68 of them, roughly ten percent. Epoch concedes the curation is subjective, and sets the difficulty bar itself — Bloom estimates that only three to five Erdős problems of this caliber had been solved by AI as of August 2026. Erdős problems have been an informal AI proving ground since early 2025, and this site covered the looser, less formal GPT-5.6-era attempts at them separately.

Lean is the referee

Natural-language proofs from AI systems are nearly impossible to verify at scale, so a solution only counts when it arrives as a formal proof in Lean, the proof assistant whose correctness checks are programmatic. Epoch built on two existing projects: Google's Formal Conjectures, which had already formalized 50 of the 68 problem statements, and Comparator, a checker from the Lean FRO designed to survive submissions that actively try to cheat. The remaining 18 statements were formalized by AI at Epoch's direction, and Epoch admits those formalizations have only passed its own initial review, not expert scrutiny — a subtly wrong statement would poison everything built on it. The protocol is strict: one attempt per problem per model, an inference budget of $300 per problem, 72 hours of working time, and no internet access, with models working from offline mathematics papers and tools like a computer algebra system. The entire harness is open source, answering an old complaint about AI math results — that nobody knew which problems were tried, by which model, at what cost.

The scorecard, and who ran it

The numbers below come from Epoch's own runs, which makes this cleaner than a vendor's self-reported table: Epoch set the budgets, ran the harness, and checked the proofs. It is still a single run of five models, and the winning column is a pre-release build, not the GPT-6 Astra shipping behind the API today.

ModelFrontierMath Erdős score
GPT-6 Astra (pre-release)3% (2 of 68)
GPT-5.6 Sol0%
GPT-5.50%
Claude Fable 5.10%
Claude Fable 50%

What the 3% actually is

Only Astra solved anything, and both results are real mathematics. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and proved problem 126 at a cost of $247 and 16 hours. Every other attempt — Astra's remaining 66, and all 68 from each rival — exhausted its budget without a verified proof. Epoch co-founder Jaime Sevilla called Astra the first model to crack a problem from this set. Separate from the benchmark, Epoch ran larger, looser attempts with the same pre-release Astra and reported further successes: disproofs of problems 1 and 74, proofs of 126, 548, and 571. Most of the remaining problems were attempted two to five times each — 172 attempts in total — and none was solved; reaching those five solutions cost over $220,000 of compute, versus roughly $20,000 for the benchmark run itself. Epoch is explicit that those attempts are not a FrontierMath Erdős score, because budgets and setups varied — they are evidence that the outcome wobbles with spending, which the fixed $300 budget is designed to control for. Astra itself had shipped days earlier; the launch coverage is separate, and the model tested here is not necessarily the model on the API.

Why the number deserves a skeptical reading

Start with formalization, which Epoch flags as an added burden. Its own example: an OpenAI model's earlier resolution of the Erdős unit-distance problem produced an 18-page natural-language proof, while the separate effort to formalize that result in Lean (github.com/plby/Erdos90) ran to 1.2 million lines of code, mostly deriving from first principles a deep theorem the natural-language paper could simply cite. A system capable of a breakthrough may still fail to formalize one, and Epoch says a model that can do both is "certainly not guaranteed."

Then contamination. No solution to any of the 68 problems was known as of August 2026, so current runs are clean, and models attempt problems without internet access. But solutions will be published and discussed, and eventually absorbed into training data; Epoch says it will monitor that and can filter out problems solved before a future model's cutoff. And the benchmark is one slice of mathematics, not a sample of it — progress here correlates with general math ability only loosely, Epoch's own framing.

The sample size deserves the bluntest statement: one attempt per problem at $300 makes these scores a lower bound on capability, not a measurement of it. Epoch's conclusion reads like the sober version of the Reddit headline — AI-driven math breakthroughs are "real but still fairly uncommon." The same week's Anthropic news sharpens the contrast: Claude formalized Fermat's Last Theorem in 11 days, a marathon of verification applied to an already-proven theorem, covered here. Solving an open problem is a different sport. FrontierMath Erdős says it can happen for $247 and 16 hours — and that most of the time, at $300 and 72 hours, it does not.

Related Articles

Scroll down

to load the next article