'FactorioBench' is trending on r/singularity — it isn't a real benchmark yet
r/singularity is hyping a 'FactorioBench' benchmark with no repo or scores. Real Astra-on-Factorio runs are underway; an open-source eval has numbers.

On September 7, a post titled "FactorioBench just dropped ;-)" appeared on r/singularity from u/BrennusSokol, drawing roughly 770 points and more than 110 comments within two days, plus crossposts to r/factorio and r/accelerate. The catch: no benchmark exists behind the title — no repository, paper, scores or methodology is linked from the post. Auto-generated trend summaries describing "sustained planning under constraints with production-collapse penalties" match nothing we could find in public.
The benchmark that isn't one
u/BrennusSokol is a regular Factorio player, not a lab, and the wink in the title reads as the joke it is. The context is real: days earlier, OpenAI's GPT-6 Astra had cleared Portal and then RimWorld, both covered here and here, and scored 99.9 percent on ARC-AGI-3 with the ARC Prize's provider-adapter harness — a result ARC Prize published on September 3. Factorio — an engineer building an ever-growing automated factory to launch a rocket — is the community's pick for the test Astra cannot pass: deterministic rules, punishing logistics, a horizon measured in hours. Commenters riffed that it is "functionally a game about designing CPUs," suggested Hearts of Iron instead, and argued brand-new games would dodge training-data contamination. Discussion, not evaluation.
Astra is already trying Factorio
The thread points at runs happening informally. On September 6, X user Mira set Astra loose on Factorio with one deliverable: "a replay file that launches the rocket in the base game." Astra designed blueprints with scripts, rushed construction robots to place them, paused the game to test designs on a headless client via Lua, and took more than 2,100 screenshots to check its work, per the post. Livestreams of Astra attempting Factorio's Space Age also circulated. Not known: whether any run has launched a rocket, and OpenAI has not commented. The same week, ARC Prize president Greg Kamradt was publicly soliciting suggestions of games Astra cannot play, Factorio among them.
The real eval already exists
Factorio agent evaluation is not new. The Factorio Learning Environment (FLE), an open-source environment from Jack Hopkins and colleagues, has run frontier models through factory-building tasks since its March 2025 paper. Its 0.3.0 results, on pre-Astra models, were blunt:
- Claude Opus 4.1 solved 16 of 24 lab-play tasks, the best showing; GPT-5 solved 15, Gemini 2.5 Pro 11, Grok 4 nine.
- No model solved any of the seven deepest tasks, from engine units to utility science packs.
- The dominant failure mode was state tracking — 97.7 percent of Claude Opus 4.1's errors.
FLE remains active — version 0.4.8 landed on September 6 — but its agents act through a Python API with the game paused, a far cry from the real-time computer-use play the meme imagines. Anyone shipping a real "FactorioBench" will have to publish the repo first. Until then, treat the hype as a wish list, not a result.


