Altman's math ladder tops out at a model nobody can test
At Dreamforce 2026 Sam Altman ranked OpenAI models by math skill, calling the ladder a framing someone gave him and capping it with an unreleased model.

On stage at Dreamforce 2026, Sam Altman ranked OpenAI's last four months of models by mathematical ability — and introduced the ranking as somebody else's idea. "A framing that someone gave," he called it, before walking up the ladder.
As he told Marc Benioff in the session posted to Salesforce's YouTube channel, GPT-5.5 was "maybe as good as an average math professor," GPT-5.6 as good as a top 1 or 2 percentile professor, Astra "a little bit better than that," and an internal model past Astra one that "can do things that the best mathematicians in the world cannot." The whole sequence, he added, covers "just the last 4, maybe 6 months."
A comparison, not a measurement
The ladder arrived without the apparatus that would make it checkable. No benchmark, no eval harness, no dated run, no scoring method. Three rungs map onto products: GPT-5.5, GPT-5.6, and Astra, the flagship OpenAI shipped on September 3. The fourth does not. By Altman's own wording it is an internal model, which puts the strongest claim about mathematics at the one point on the ladder no outsider can reach. OpenAI's disclosure stream has described unreleased internal models before, so the shape of the claim is familiar; what remains missing is anything a reader could open. He offered a parallel assurance on safety, saying he was "very confident" the industry could keep "alignment and safety and monitoring way ahead of capabilities," and citing "the Hugging Face accident" as a sign it no longer "takes as much imagination as it used to to imagine how this could go wrong."
The checkable contrast
Earlier in the same conversation, asked why AI fear had gone mainstream, Altman said models had moved in about three years from ones that "could barely carry on a conversation" to ones that "can prove Millennium Prize problems." Six of the seven Clay Mathematics problems remain open, and the Riemann hypothesis is exactly the kind of statement nobody has settled.
Public results in the same period were narrower and more specific: Astra's first solutions on Erdős problems carried named problems, and OpenAI's September 8 Navier-Stokes post went out as a proposed solution, which is not the same thing as a verified one. A r/singularity thread carrying the clip was sitting at roughly 740 points and more than 310 comments a day later, though such numbers are approximate.
What would settle it
A named model. A published evaluation, with the problems listed and the scoring visible. And a third party able to rerun it — the same standard the labs apply to themselves when they publish a benchmark table. Until then, this is a comparison someone handed Altman on a stage run by a partner company.


