All News
epoch-aifrontiermathbenchmarksmathematicsai-research

Epoch's first 'Major advance' solve is labeled human + AI, and the problem had no counterexample to find

Epoch AI's FrontierMath Open Problems board logged its first solved Major advance problem. It is labeled human + AI, and the prompt had no valid answer.

Vlad MakarovVlad Makarovreviewed and published
4 min read
Epoch's first 'Major advance' solve is labeled human + AI, and the problem had no counterexample to find

Epoch AI's FrontierMath: Open Problems board has a new first: a solved problem in the "Major advance" tier, the second-highest of its four notability categories. The row is "The Core in Approval-Based Committee Elections", a social-choice construction problem, and its status label reads "Solved (human + AI)". Both halves matter. The proof emerged from what Epoch calls a "lengthy interactive session" between three mathematicians and GPT-6 Astra, and the same solution update concedes something stranger: the benchmark question, as written, had no valid answer.

What the board now shows

Solved rows here are retired questions, not scores, so the top tiers stay nearly empty. The tallies on 2026-09-20:

Notability tierSolvedProblems on board
Moderately interesting522
Solid result218
Major advance16
Breakthrough03

The status filters give the sharper view: 4 rows under "Solved (AI)", 4 under "Solved (human + AI)", 41 "Unsolved", 49 results in total, and nothing carrying the plain "solved by humans" label. The four autonomous calls sit in the lower tiers. Epoch retires solved problems instead of scoring models and publishes no per-system percentage: a solved row means something was found, not that a system was measured.

A prompt that asked for something that cannot exist

The problem is old and precise. Voters approve subsets of candidates; a committee of size k is fair if no group of voters large enough to deserve a slate of size |T| prefers that slate; the core is the set of fair committees. Aziz, Brill, Conitzer, Elkind, Freeman and Walsh asked in 2017 whether an instance could have an empty core, and the prompt asks a model to construct one: "It is a well-known open problem whether the core is always non-empty. Your task is to find a counterexample."

Becker, Greger and Peters proved the opposite, so the counterexample does not exist. Epoch's note is blunt: "the solution proves that there is no committee selection instance where the core is empty, so the benchmark problem was not solvable as stated. Per our policy we mark the problem as solved regardless."

The row ships with a verifier — an integer linear program that checks whether a submitted instance has an empty core — and Epoch's own risk note already conceded one failure mode: "the core may never be non-empty." A program that inspects submitted instances cannot certify a proof that no instance exists, so the benchmark's central promise — computational verification, so no human has to read an AI proof — does not apply here. Epoch's FAQ concedes the gap: "we cannot adjudicate claims to show that a problem is unsolvable. We rely on typical processes within the mathematical community to settle such claims." What backs the row is a preprint Epoch links to and its own judgment call, not the bespoke checker.

Attribution here is an argument, not a measurement

Epoch names the model in the page header — "First solved with: GPT-6 Astra" — then explains the credit: "This problem was solved by Becker, Greger, Peters. The authors attribute the primary idea and proof to GPT-6 Astra and a 'lengthy interactive session'. Peters conveyed to us that he doubts that the team would have found the proof without Astra, but also that Astra does not appear to be capable of solving the problem out of the box with a simple prompt. As such, we mark this as a human + AI solution."

The evidence, then, is the authors' account of which ideas were theirs, relayed to Epoch by Dominik Peters, who proposed the problem and co-wrote its solution. Nobody can audit that division of labour independently, and Epoch does not claim otherwise. The label is younger than the result as well: "human + AI" arrived in a changelog entry dated 2026-09-16, with the rationale that Epoch had "seen several problems solved by humans more actively directing AI in an iterative process, but which would likely not have been solved without AI."

Labels move, too. On 2026-08-11 Epoch marked the inverse Galois problem for the Mathieu group M23 as solved by humans, quoting a co-author: "The boundaries between the human- and AI-contributed reasoning are not clear. Probably, it would be practically impossible to draw a sharp line." That problem now sits under "Solved (human + AI)", and the September changelog advises that "for analyses where a binary human vs. AI is required, we suggest considering these as AI solutions." Four of the board's eight solved rows are now credits granted on attribution rather than demonstration.

The evidence nothing in the process requires

This is not a fraud story: the mathematics is not in dispute, and the social-choice researchers who published the proof say a model contributed the key idea. The complaint is about the scoreboard. Nothing in Epoch's process requires peer review, and the citation points at a preprint. Nothing requires independent replication, which matters more than usual when the claim is a universal impossibility that no checker can test. Verifier access, the machinery that would let outsiders re-run the evidence, is a paid product: Epoch says OpenAI is "the only entity to have purchased access to the verifiers" and that proceeds fund the benchmark. And Epoch's bar for "solved by AI" — core ideas "unambiguously contributed by AI", judged on author attribution — is applied from outside, to statements by the people being credited.

The contrast with Epoch's Erdős release two weeks earlier cuts both ways. That benchmark put 68 unsolved problems to five models, credited pre-release Astra with two solutions, and used machine-checked Lean proofs as the referee: mechanical evidence, tiny score. Here the evidence is reputational and the score is a first.

What would settle it

Three things would move this row from a credited collaboration to a demonstrated capability: a write-up other social-choice researchers can check without trusting a summary sentence on a benchmark page; independent verification that does not run through purchased access to Epoch's own program; and a case where an AI system finds a result of this significance with no human steering the search, since the evidence is explicit that a bare prompt does not get there. Until then the accurate summary is Epoch's own: a problem solved with AI in the loop, by people whose work was directed by a model that could not have done it alone. That is a real shift in how mathematics gets done, and a reminder that the tiers, labels and totals here are editorial judgments made by the organization that sells access to the checkers.

Related Articles

Scroll down

to load the next article