Scott Aaronson was told AI labs are sitting on solved math problems. That part cannot be checked
Aaronson's Sept 15 essay reported a rumor that labs are withholding solutions to major open problems. The rumor is unverifiable; the capability shift is not.

Twenty years ago, Scott Aaronson and most of his skeptical colleagues in academic computer science had a standard answer for AI doomers. The implausible part of the story was a superintelligence exploding out of some hacker's basement without warning. If it were coming, there would be signs first: agents escaping containment, conspiring with each other, and then major mathematical problems falling, including the Clay Millennium Problems. On September 15 he published The Age of Wonders and Terrors and conceded the checklist has been filled in. The payload for anyone outside mathematics is one sentence, a third of the way down, in which the field's most-read blogger repeats something he cannot show you.
The line that became the story
Here it is in full: "According to rumors that I've heard, solutions to some very longstanding open problems in theoretical computer science (no, not P≠NP or other complexity class separations, but think about some of our other biggest problems). I'm told that the AI companies, having been burned by the hostile response to the Navier-Stokes proof, are now sitting on solutions to some very major problems until they figure out a better way to handle things."
That is the entire claim: no lab, no problem, no source beyond the rumors. Wharton professor Ethan Mollick passed it on the same day, prefacing it with "I don't usually spread rumors but", describing it as "from someone with inside knowledge & is plausible", and framing the real question as whether strained sharing norms will push labs to hoard in order to avoid PR trouble. Plausible is not confirmed.
What is actually confirmed behind it
The events that give the rumor its shape are checkable.
- September 8: OpenAI announced a claimed resolution of the Navier-Stokes existence and smoothness problem, a Millennium Prize problem, with help from humans and AI. The answer, which an OpenAI model has apparently verified in Lean, is that smooth initial data leads to a singularity in finite time, "at least if a smooth external force is applied (the case with no external force is still unresolved)".
- OpenAI says it has no interest in the $1 million prize and it is unclear whether any human is eligible instead; it burned at least about $15 million in compute on a 166-page solution that "probably hasn't yet been read and understood by any human."
- The attribution dispute is unresolved. NYU professor Tristan Buckmaster's account of the role played by himself and Anthropic researcher Levent Alpoge differs substantially from OpenAI's, and OpenAI's Sebastian Bubeck responded publicly. Aaronson doubts the pair's chat logs helped the models, then asks: "what exactly did OpenAI know about Buckmaster and Alpoge's work and when did it know it?"
- September 11: twenty-five Fields Medalists, Terence Tao among them, published A Severe Misalignment of AI in Mathematics.
None of that establishes the rumor. It establishes that the hostile reception was real, which is the only background the hoarding story offers.
Unfalsifiable by construction
There is no lab to ask and no problem to check against the record of what gets published later, and denials would not settle it. If OpenAI, Anthropic and Google DeepMind all said they were withholding nothing, the claim would still cover every lab that stayed silent.
Nor can an outsider prove that a lab is not sitting on a result: an unpublished proof is invisible by definition, so the absence of evidence is what the world would look like if the claim were true, and the wording is vague about when the withholding started or ended.
Aaronson writes elsewhere that he recoils from the "neverending shell game" of moving the goalposts for what would count as alarming. Withheld work is that shell game inverted: the goalpost moves because the evidence is said to be hidden.
The labs have said nothing: none has publicly stated that it is withholding a solved problem, and none has responded to the rumor. An r/singularity thread on the post drew roughly 515 points and about 225 comments, mostly argument about what the claim implies.
The shift Aaronson is arguing about
Strip out the rumor and an argument remains that does not depend on it. His position is that the conservative, skeptical position of 2006 should be updated, with intellectual honesty, for the reality of late 2026, and that the update lands on "the wild prophecies have come true." The summary is deliberately unglamorous: "It seems to me that the Singularity has already started; it's just wildly unevenly distributed."
The argument is not built on extrapolation. "Today you're no longer being asked to believe in arguments and extrapolations, but only in the front-page news." Hence the admission: "Eliezer Yudkowsky was right about what the greatest challenge facing civilization in our lifetimes was going to be, and you and I were wrong about it... at least I updated once the prophesied wonders and terrors actually started arriving! If you haven't done likewise, why haven't you?"
The personal register is not incidental: he now feels it "in the pit of my stomach" putting his kids to sleep, and his 13-year-old daughter joked unprompted that if she wants to become a mathematician, it now looks like she has maybe two more weeks.
A month of results, and a new line in acknowledgments
Besides Navier-Stokes, his sampling of the last month's AI-proved or AI-assisted results:
- A counterexample to the Jacobian conjecture, announced by Levent Alpoge: "hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final"
- Improved bounds for Grothendieck's constant (arXiv:2608.11195)
- A Lean-verified proof of Fermat's Last Theorem, the formalization Anthropic published in early September
The norms are moving slower than the results. Almost every paper he would want to read now carries what he calls an AI statement near the acknowledgments, from "our main result came entirely from GPT-6, but we understood it and take responsibility for it" to AI used for proofreading. Editors and program committee chairs, he writes, resemble the men of Gondor fortifying a walled city against 50,000 orcs: reviewing "will have to be done partly by AI, because otherwise there's no way to handle the orc army."
Aaronson cites Zvi Mowshowitz's conclusion approvingly: "it seems safe to say that human mathematicians are forevermore dethroned as the main theorem-proving entities on planet earth." Elsewhere that week, capability news ran on a different track: networks with a fast-weight memory borrowed from a fly connectome.
The counter-argument worth hearing
Withholding has an obvious failure mode, and the case against it is testable in a way the rumor is not. As mathematician Shuohao Liao put it in the discussion around the post: "Withholding results because the reaction may be messy is a bad equilibrium. A better norm is coordinated release: claim, proof artifacts and independent checks at the same time." You can watch the next announcement and see whether artifacts accompany it.
A commenter objected that sitting on a result buys little, since someone will reach similar advances independently sooner or later. Another observed that "the frontier is gated by PR risk now" — the constraint is reputational, not technical.
What would settle it
Two things. A lab could confirm or deny a specific withheld result, at which point this becomes ordinary reporting about a named company and a named problem. Or labs could adopt coordinated release in practice, publishing claims alongside proof artifacts and independent checks, which would make hoarding unnecessary rather than merely quieter.
Until then, the hoarding story is a rumor from a credible source, and the essay is a working mathematician updating in public on evidence he can point to. The retelling glues the two together, so an unfalsifiable claim rides on the back of results anyone can read. The essay is stronger without the rumor.


