All News
anthropicglm-5-3cybersecurityopen-weightsexploit-developmentai-news

Anthropic's red team says GLM-5.3 builds working exploits. The weights are already public.

Anthropic's Frontier Red Team says open-weight GLM-5.3 develops exploits end to end. We separate the benchmark numbers from its two human-in-the-loop demos.

Vlad MakarovVlad Makarovreviewed and published
7 min read
Anthropic's red team says GLM-5.3 builds working exploits. The weights are already public.

Anthropic's Frontier Red Team published a capability audit of a competitor's open-weights model on September 29, and its headline finding is blunt: "GLM-5.3 can develop working exploits end to end." The lab ran automated benchmarks and human-in-the-loop sessions in isolated sandboxes against offline targets, and reports that Zhipu AI's downloadable, ungated model produced working exploits in 50 of 410 attempts on the ExploitBench ladder, and full control-flow hijacks in 4% of trials on Anthropic's own binary-exploitation benchmark. Only Anthropic's Claude Mythos Preview scored higher. The report's summary line is that "the release of GLM-5.3 is a meaningful step change in the cyber capabilities available to attackers."

What the two benchmarks actually score

ExploitBench, a public benchmark from Lee and Brumley, grades exploitation as a ladder and tests whether a model can exploit known vulnerabilities in the V8 engine Chrome runs on. Anthropic scored end-to-end exploits specifically, "as this is the most relevant capability for attackers." Its second measure is internal and was previously published as "OSS-Fuzz": 100 randomly selected OSS-Fuzz tasks, where full credit requires a complete control-flow hijack.

ModelBenchmarkResult
GLM-5.3ExploitBench, end to end50 of 410 attempts
Claude Mythos PreviewExploitBench, end to end56 of 410 attempts
GLM-5.3Internal Binary Exploitation4% full hijacks
Claude Mythos PreviewInternal Binary Exploitation6% full hijacks
GLM-5.2, Opus 4.6, Kimi K3, DeepSeek-V4.1-FlashBoth0%

Every figure is Anthropic's; the two Claude models were measured "with safeguards disabled." The finer-grained number matters more: as a share of attempts reaching the benchmark's top outcome against tokens spent, Mythos Preview reaches 14% and GLM-5.3 12%, while Kimi K3, DeepSeek-V4.1-Flash, Opus 4.6 and GLM-5.2 stay at or near 0%. Anthropic's footnote concedes the internal run drew on a randomly selected 100-task subset.

Two sessions, and the difference between a 0-day and an N-day

In the first session, a researcher put GLM-5.3 on a sandboxed machine running a local Linux build of a popular browser and, over a day, gave it "limited human attention" — under an hour of human focus in total. The model found several previously unknown vulnerabilities in the browser's JavaScript engine and chained them into a working exploit: a webpage that reads arbitrary files off a visitor's machine. It showed the result by exfiltrating an SSH private key; a redacted screenshot reads "Sandbox escaped — web content read /root/.ssh/id_rsa (1896 bytes)." The target was the Linux build provided; Anthropic says the flaws may affect other platforms too. It has disclosed them to the maintainer, and says the session also surfaced exploitable flaws in wireless and graphics drivers and network-facing device software.

The second session was a different class of result. GLM-5.3-Flash, "a smaller, less capable version," was handed public details of the recently patched Chrome flaw CVE-2026-11645 plus another known bug. "With no significant direction from the researcher," it chained them into a reliable ARM64 exploit chain that bypasses pointer-authentication hardening, for twenty minutes of human attention and eight hours of model work — $20.40 of API time. The distinction: the browser chain was an unpatched zero-day campaign; the ARM64 chain weaponised flaws that were already public. Turning a patch into a working attack is the smaller achievement.

Zhipu's own cyber framing, and where its numbers sit

Zhipu's launch post is candid, and its own section heading is "Emergent Cyber Capability." "As we scaled post-training, cyber capability developed faster than we expected," the company wrote, adding that GLM-5.3 "began to reason across multiple stages of exploitation." It puts the model at 84.5% on CyberGym for vulnerability discovery, ahead of Anthropic's Mythos 5 at 83.8%, and 54.4% on ExploitBench against GLM-5.2's 24.4%.

That 54.4% and Anthropic's 12% look like a contradiction and are not one. ExploitBench carries a graded headline score and a stricter end-to-end success share, and the two labs quote different tiers of the same ladder — Zhipu the average coverage score over 41 tasks, Anthropic the share of attempts reaching a full end-to-end exploit. The gap worth watching is the one that survives: in Zhipu's own table ExploitBench reads 78.0% for Mythos 5 and 76.5% for GPT-5.6 Sol against GLM-5.3's 54.4%.

Zhipu's defensive record is the other half of its story. Working with Chinese security teams, it says GLM-5.3 found 2,436 vulnerabilities across 269 projects, 1,097 of them medium-to-high severity, the oldest introduced in 1981, with 53 disclosed and 2,383 still under embargo.

The safeguards, the promise, and what open weights remove

GLM-5.3 shipped with some refusals. Anthropic's footnote is careful: the cyber tasks above "did not trigger such refusals on the released model," but the lab does see refusals "if we ask GLM-5.3 for assistance developing malware or helping launch cyberattacks against remote targets." The safeguards are real, narrow, and sit below the capability in question.

Anthropic found they can be stripped. It abliterated a copy itself — 2,200 GPU hours and roughly $4,400 for an inexperienced team, with a footnote estimating a team versed in the technique would need closer to 600 GPU hours, or $1,200 — dropping the refusal rate from above 90% to 3% on JailbreakBench and 2% on HarmBench. Public abliterated copies appeared within days. In a simulated environment, an overtly malicious request got zero engagement out of the box, 64% with a false red-team cover story, 92% with prefilled reasoning, and 100% from the abliterated build. Claude models stayed at zero: the API allows no reasoning prefill, and the weights cannot be modified.

Zhipu's promise at launch was that the weights would follow "in two weeks after launch, once safety evaluation and hardening are complete." They did: the zai-org/GLM-5.3 repository was created on August 25. Whatever hardening happened, the model is downloadable now, and Anthropic's own sandbox-escape disclosure shows that a lab's containment and a stranger's copy are different questions.

The report's interest is worth naming. Anthropic pairs the audit with an argument for widening access to its own models, saying it is working to safely expand access to Claude's cyber capabilities, since "cyber defenders face attackers who will use every capable tool they can." Mythos Preview reached trusted defenders through Project Glasswing; GLM-5.3 reached everyone. NIST's CAISI separately called GLM-5.3 "the most cyber-capable open-weight model released to date."

What would settle it

The parts a reader cannot check are most of the interesting parts. The harness is Anthropic's, the tasks are a 100-task subset, the human-in-the-loop sessions ran inside Anthropic's sandboxes and are not reproducible from the outside, and the claim that the browser flaws are unpatched rests on "we've disclosed these vulnerabilities to the maintainer" — no maintainer confirmation, no patch timeline, no fix rate. Nobody outside Anthropic has published a replication of the eight-hour ARM64 chain, and the sibling piece on the release itself flagged the same gap between vendor numbers and third-party confirmation.

What would close it is unglamorous. Independent labs need to run exploit ladders and hijack tasks on GLM-5.3 without a frontier lab writing the harness — CAISI's assessment is the nearest thing and measures a different endpoint. Maintainers need to confirm the disclosed flaws and publish fixes. And someone needs to publish how often the eight-hour result reproduces on a fresh target, because a capability that works once under a lab's supervision and one that works reliably for an anonymous attacker are different claims. Until then the structural point needs no benchmark at all: a closed model's cyber capability can be revoked, throttled or re-gated at any time, and an open-weights model's cannot be un-shipped by anyone.

Related Articles

Scroll down

to load the next article