An AI built its own graphene simulator, then broke its own design rule
An MIT lab's AI agent wrote a graphene simulator from scratch, ran it for days and found a 6.6x strength spread. The caveat behind its viral 25% claim.

On September 28, Markus J. Buehler, an MIT professor of engineering who leads the Laboratory for Atomistic and Molecular Mechanics, published a long first-person essay on X about an AI agent that wrote its own atom-by-atom simulator for graphene and then ran virtual experiments with it for several days. The result the essay leads with is a spread of more than sixfold in density-normalized strength across graphene architectures. Two days later a shorter follow-up post drew far more attention with a different number: designs "about 25% stronger." Those two figures are not the same claim, and the gap between them is where the useful reporting sits.
What the agent was asked to build
The experiment started with a prompt and five images suggesting leaf veins, nested networks, disordered fibers, crossing struts and rings around a void. The agent had to infer a design language from the images, construct the atomistic instrument "from scratch," and commit to predictions before running the simulations that would judge them. In practice that meant implementing energy and force calculations in PyTorch, plus geometry generators, neighbor lists, relaxation and loading routines and an experiment database — thousands of lines of code, checked against reference calculations. A discrepancy in that comparison was traced to the rounding of numerical constants; correcting it reduced an energy difference from about 10⁻⁸ to 10⁻¹³ electronvolts per atom. The summary post counts twenty validation tests behind that. The reasoning phase ran for several days and covered hundreds of discovery simulations across many architecture families. The output, in Buehler's account, is a library of hundreds of thousands of atomically specified graphene structures. His group has published automated-discovery work before, including SciAgents and a physics-aware multiagent system for alloy design in 2025, so this reads as the continuation of a program rather than a one-off.
The result the essay leads with
Among dozens of design families with relative densities between 0.77 and 0.83, density-normalized strength spanned a factor of 6.6. Nearly the same amount of carbon carried sharply different loads because the atoms were organized into different paths for force. The agent then proposed a rule — the solid fraction at the narrowest load-bearing section, multiplied by the strength of the relevant ligament — and the rule's own test broke it. A slit array tilted by 20 degrees was predicted to be much stronger than the simulations delivered. Chasing that miss moved the agent to the bridges between neighbouring slit tips and to a staircase-like fracture path, and further tests separated three regimes, in which neighbouring slit tips link, intervening ligaments rotate, or short bridges bend. Elsewhere, an apparent optimum in a disordered network disappeared under seed replication. An unsupervised embedding sorts the surviving designs by geometric similarity, and the families it exposes run from regular perforations and angled cuts to branching networks, curved reinforcements, graded textures and hybrid domains.
The 25% figure is real, and narrower than the headline
The viral number is not invented. It traces to the follow-up post, and it arrives with conditions attached. There, Buehler writes that at the original scale much of hierarchy's apparent strength advantage "can be explained by alignment," and that only "with greater separation between structural levels ... selected hierarchical designs became about 25% stronger than same-mass single-level controls." Every qualifier is load-bearing: selected designs, larger scale separation, matched mass, and a simulation. The repost that spread the claim reduced all of it to "found graphene designs ~25% stronger at the same mass." A commenter under that thread asked the fair question, whether that was simulated 25% stronger or verified in a lab 25% stronger. It is the simulated kind. The summary post gathered roughly 1,300 likes against 235 on the research essay as of this writing, which is part of how a conditioned sub-result became the headline. The main essay never mentions 25% at all; its headline finding is the 6.6x span.
The loop is the real claim
Strip away the materials framing and the interesting object is a sequence: build the instrument, state a prediction, run the test, accept the miss, and trace it to a mechanism. The failed tilted-slit prediction and the 10⁻¹³ rounding fix are the same move at different levels. A commenter writing as @WorldianAI put it plainly: "The part that matters is the loop. It built the instrument, guessed, got a tilted-slit prediction wrong, and used the miss to find a mechanism." That is a more defensible claim than any strength figure, because it is about process, and the process is legible in the record even where the physics is contested.
Where the loop stops
Everything here is simulation. No sample was made, no membrane was pulled, and no measured strength appears anywhere in the material. The essay describes the work as a preprint; its reference list points to ChemRxiv under a title about "models building models," and no matching preprint page or DOI surfaced when I looked, so Buehler's own posts are the only checkable sources. The deeper limit sits inside the instrument. The agent can revise its mechanisms when a simulation falsifies a prediction, but the simulation itself is fixed by the chosen interatomic potential — the reference list points to Brenner's reactive bond-order model and later screened variants. A commenter, @yinshuangx74130, drew the line: "can the agent ever discover that its microscopic model itself is wrong, rather than only revising explanations within that model?" Another asked how anyone verifies the simulator has no subtle bug, noting that "there is no Lean to prove correctness in physics."
Not every objection was technical. One reader on the summary post argued the whole exercise was obvious: carbon forms four bonds, the bond lengths and angles are known, and arranging known sheets for strength is, in his words, "something obvious to any engineer." That reading understates how quickly the combinatorial space of atomic arrangements outgrows intuition, but it is a fair reminder that the novelty on offer is the agent's loop, not the chemistry of carbon.
What would settle it
Three things, in rough order. A laboratory synthesis and mechanical measurement of one candidate architecture, which would test the simulator against reality rather than against itself. Peer review of the preprint, which currently cannot be found. And independent replication of the instrument by a group that does not work with Buehler, since an agent-written simulator is the part most in need of outside eyes. The essay notes that the instrument and its evidence persist beyond the initial investigation, which is exactly what would make such replication tractable. The closest comparison this month, Anthropic's claim that one of its models spotted a CRISPR-like repeat array in viral DNA, at least carried bench experiments run by human scientists before anyone wrote a headline. For scale, the only graphene strength figure that has survived that treatment is a measured one: in 2008, Changgu Lee and colleagues reported an intrinsic strength of 130 gigapascals for free-standing monolayer graphene, obtained by nanoindentation rather than by a model.
The smaller claim
What exists today is a design-principles library and a demonstrated loop, not a material. Buehler's own essay is careful about that; the summary post and the reposts less so. The sixfold spread is the finding. The 25% is a conditional sub-result about a subset of designs at matched mass, and it is the number most likely to outlive every caveat attached to it. Treat it as an invitation to ask what was measured, and the answer is nothing yet.


