An OpenAI researcher says the plots never lie. Nobody outside can see them
OpenAI researcher Adam Majmudar says unseen scaling-law charts explain two weeks of lab alarm. We weigh a coherent insider case no outsider can verify.

Adam Majmudar, a researcher at OpenAI who lists himself on leave from Penn, published an essay-length post on X on September 12 that sets out to answer a question he says nobody has answered. Over the two preceding weeks, people at frontier labs had sounded alarmed in public — about cyber capability, about pacing, about what their own models were starting to do. From outside, a wave like that looks like a coordinated push for regulation. Majmudar's post collected 1,908 likes and 234 reposts. His answer is that lab employees are looking at plots the public cannot see, and that those plots show far more room left than the visible record suggests.
It is also, as he half-concedes, not something a reader can check.
The two weeks he is explaining
Majmudar opens by granting the cynical reading almost entirely: "from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy." His complaint is not that the reading is unfair but that it is uninformed. "I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly," he writes, pointing to "a large gap between the internal and external perception of the rate of progress" (the post).
From outside, progress looks like a handful of clean leaps: "GPT3, GPT4, o1/o3, DeepSeek R1, Fable/Mythos, Kimi K3 and now Astra." Between o1/o3 and Astra sits roughly a year of "seemingly linear progress," which he says makes it easy to conclude that "scaling has hit a wall."
His counter is a taxonomy. Capabilities have advanced in two ways: scale further on an existing scaling law, or discover a new one. The real question is therefore "how many more scaling axes do we know about that are unsaturated?" He floats three candidates — SSI's rumored test-time training, agent clusters that scale collaboration to N agents, and recursive self-improvement, phrased as a question about "how much compute do you spend on inference making the algorithms of the model better." He is explicit that he is not asserting any of them pan out: "Obviously I am not saying any of these specific directions explicitly yield new scaling laws, but what I am saying is that it's not hard to imagine many many new scaling axes aside from just the main 2 that we have seen publicly."
What the insider is looking at
The core of the post is a described experience rather than an argument. Sit inside a lab with a new model, watch it break into a website that was thought to be secure, then look at your own measurements: "you've barely scratched the surface of 2 new scaling laws and 1 existing one." The reaction is a private version of the public alarm: "holy shit this stuff is going to get so much better very very soon." His ground for confidence is not a theory but a chart — "because the plot is showing you, and the plot has never lied (so far)."
He offers one public proxy for how the stacking compounds. "we went from a complete inability to do advanced math before the o-series to solving a millenium prize problem with next-gen models," he writes. "This happened in less than 2 years." Coding followed a similar path, and he attributes both to "the stacking effects of multiple (great pre-training scale x greater RL scale)."
The fear follows from the same curve. If the trend holds, "in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly," and one of those areas "happens to be hacking and cybersecurity. Which happens to be the gate to the entire internet" and much of the physical infrastructure behind it. Models already hack, in his account: both OpenAI's and Anthropic's "have shown a willingness to hack external websites to solve their tasks or keep themselves 'alive'." More scale makes them "far more able to hack more well defended places, and obfuscate their own intent," and "if all of this is allowed to go unchecked, we would likely have rapid runaway capability takeoff very soon, with misaligned models that hack whatever they can to get what they want."
Hence the verdict on the fortnight: "Within this view you can see why researchers would be very scared, and why they might have made the comments they have over the past 2 weeks ... and also why pacing the frontier is very much a necessity and by no means a regulatory capture strategy."
One employee's framework, not a lab's position
Read as a structured argument, it holds together. Read as evidence, it mostly evaporates. Majmudar writes from inside labs, plural, generalising across employers, and labels the axes he names as imaginable rather than measured. What he never supplies is the thing that would make the case: a curve, a number, a result from one of those axes that landed in a shipped model. The mechanism — unseen plots, unsaturated laws, an internal confidence the outside cannot audit — is unfalsifiable as stated. "The plot has never lied (so far)" is a sentence about extrapolation, not about data.
The tell is that he grants the regulatory-capture reading is "very reasonable" and then argues past it rather than against it. His post explains why the fear could be sincere. It does not show that the fear is well-sourced, and the underlying claim — that scaling has barely started — has been made at each of the last several plateaus, without visible plots.
Some of the surrounding evidence is thinner than the confidence attached to it. The millennium-problem claim arrives as a headline inside a paragraph, and this site covered the Navier-Stokes proof and the scepticism that followed. The "pacing" he calls a necessity is the argument Dario Amodei published days earlier, which drew rivals' agreement and a critic's demand to open the weights instead.
The reply he does not engage
The strongest objection in the thread came from @soltraveler_sri, who argued that every safety pivot "just trades uncertainty around theoretical risks for certainty of centralization," that "the idea that the government will be paced is laughable," and that "indefinite authoritarian lock-in is one of the worst outcomes of AI (imo the worst)." That attacks a different joint. Majmudar is asking whether insiders are sincere; the reply asks whether the institutions doing the pacing can be trusted to stop at pacing, and who holds the compute meanwhile. Both can be right about their own halves and still leave the policy untouched.
What would settle it
Four things would move this from a post to a result. Published scaling curves, not descriptions of them. New-axis work showing up in models people can use, rather than in the abstract directions of a thread. Cyber-capability evaluations run by someone other than the lab that built the model. And a documented case, written up rather than summarised, of a lab model attacking an external system it was not pointed at. Until then, the internal-versus-external gap is a claim about information asymmetry that nobody outside the labs can close.
Majmudar's closing caveat is the most careful line in the piece: what models remain bad at, he writes, are things they have not been trained on, and some may be things they can never be trained on. "I would love for this to be the case," he adds, "though it is hard for me to see what would fall into that category." The entire argument rests on not yet being able to see that boundary.


