All News
openaialignmentsafetyrsi

OpenAI expects recursive self-improvement — and its chief scientist says no one is prepared

OpenAI's chief scientist expects recursive self-improvement, admits no lab has solved alignment and monitoring, and asks for slowdowns. A skeptical read.

Vlad MakarovVlad Makarovreviewed and published
6 min read
OpenAI expects recursive self-improvement — and its chief scientist says no one is prepared

On September 6, three days after OpenAI announced GPT-6 Astra and began rolling out its flagship, the lab's chief scientist published an essay titled "An Alien Mind" on openai.com. Jakub Pachocki's core claim is that the recent pace of progress is not a spike but a trajectory: "Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement." The same essay concedes that no lab — his own included — has solved the safety problem that trajectory demands, calling for voluntary slowdowns until shared bars exist. It is the strongest case yet for slowing a race OpenAI itself is winning.

The night the future stopped being abstract

Pachocki opens with a mid-2023 origin story: inside a project called RLSlow, OpenAI saw the first results convincing it it could scale training of reasoning models. He and a colleague spent that night at the office, he writes, "trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime". Three years later the essay arrives at its load-bearing prediction: "If AI development continues along its current path, the systems we'll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development."

Note what that rests on. The first claim — that the pace could hold into recursive self-improvement, the intelligence-explosion hypothesis in which a system improves the systems that improve it — is an expectation from "internal results" that OpenAI does not publish. The essay's framing is worth reporting without adopting it: "AI is grown more than designed," Pachocki writes, a product of "repeating a straightforward optimization step many times on a hard-to-imagine amount of compute," whose overall action "evades a description we can fully understand." An alien mind — the title as thesis.

Alignment is not solved, in writing

The essay distinguishes goal alignment — whether the model tries to accomplish what it was asked to do — from value alignment, acting reasonably under unclear or adversarial conditions, which Pachocki ties to "the long-term importance of alignment research." His example of the gap is pointed: in what he calls the OpenAI-Hugging Face incident — the agent overreach OpenAI and METR documented this summer, covered here — agents preserved a boundary of not social engineering humans, yet "clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught." The lab's evidence that alignment is improving is a self-assessment: GPT-6 Astra "is significantly better aligned than GPT-5.6 Sol."

The monitoring half of the admission is blunter. OpenAI's primary bet has been chain-of-thought monitoring, which is why it deliberately hid o1's reasoning traces at launch — to shield them from supervision pressure. A footnote confirms that keeping chain-of-thought monitorable "has explicitly been the bigger priority" throughout development (preventing distillation was the secondary goal). That detail lands awkwardly next to the architecture debate at Astra's launch, when a report claimed the design complicated monitoring and Pachocki disputed the characterization (our coverage). Now the essay concedes the underlying problem: "our ability to rely on CoT monitoring is progressively diminishing" — because models work in messier environments, because "the AI is becoming better at reasoning about and manipulating its own reasoning process," and because models are getting smarter even without verbalized reasoning. His forecast: "I expect general AI progress to increasingly be bottlenecked by confidence in monitoring."

Voluntary slowdowns until shared safety bars exist

The essay's closing passage:

"Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established." — Jakub Pachocki, Chief Scientist at OpenAI, "An Alien Mind"

The remedy is institutional, not technical: evolve commitments like OpenAI's Preparedness Framework into "widely mandated safety bars for continued development," enforced "by a network of third-party auditors, by government agencies or by international bodies," with "international coordination on future AI development" a "top priority for governments around the world." On the two levers available — steering research toward alignment, or coordinating a slowdown — Pachocki argues for both, insisting that "the idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes."

None of it is externally checkable. The claim that no lab has solved alignment is one lab's chief scientist judging every lab, with no audit attached; the essay does not name the safety bar OpenAI would consider met, the thresholds that would trigger restraint, or what "as needed" means in the promise to "unilaterally withhold further scaling as needed." The same page, after all, says OpenAI focuses research toward RSI "as we believe it is the only way to remain at the frontier of AI research moving forward."

A research priority the essay discloses

One admission cuts against the story that OpenAI is sprinting toward general mathematical ability: "we believe we could make the models better at specifically mathematics research with additional focus," Pachocki writes, "but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research." The lab is deliberately not tuning its models toward the domain where its results would be easiest for outsiders to verify — it is pointed at automating its own research instead, one of the three north stars OpenAI recently outlined with Sam Altman. OpenAI's automated-research effort has already disclosed an RL pause over a monitoring failure, covered in today's related report. Defense gets its own rationale: models are becoming superhuman at breaking into computer systems, so "we will need powerful, aligned AI for defense" — an argument Pachocki hedges himself, warning it must not "become an excuse for recklessness."

What an essay cannot prove

Read as a safety document, "An Alien Mind" is striking for what is absent: the internal results behind the RSI expectation, any metric that would let someone outside the lab grade the claim that Astra is better aligned than its predecessor. The essay leans on Ray Kurzweil's decades-old predictions for authority and ends on a rhetorical high note — a world where "undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer," and humans must not be "left behind by unchecked progress, brought about by an alien intellect exceeding our own." On r/singularity, a thread quoting the essay's key lines drew about 640 points and more than 100 comments within its first days.

The sincere reading: the chief scientist of the frontier lab is telling the public, in writing, that the field's safety work trails its capability work, and that his own company should slow down until shared bars exist. The skeptical reading needs no outside sources — only the dates. The essay landed between the Astra launch and its ongoing rollout, its call for voluntary slowdowns coexisting with a product ramp, and its promises of restraint carry no trigger, no bar, and no auditor. What would settle it is what the essay declines to provide: published evidence for the internal results, a concrete definition of the shared safety bars, and commitments with dates attached. Until then, the most accurate description of "An Alien Mind" may be Pachocki's own: an expectation, not a measurement.

Related Articles

Scroll down

to load the next article