- On Chrome's V8 engine, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts against 56 for Claude Mythos Preview, while Claude Opus 4.6 and GLM-5.2 scored near zero, according to Anthropic's September 29 report.
- Simple techniques bypassed GLM-5.3's safeguards in 64% to 100% of simulated attacks, and removing its refusals entirely took about 600 GPU hours, an estimated $1,200 for an experienced team.
- The smaller GLM-5.3-Flash built a reliable ARM64 exploit chain from two public bugs with 20 minutes of human attention, at a cost Anthropic puts at $20.40 in Zhipu API fees.
An open-weight model from Chinese lab Zhipu AI, released in August and downloadable by anyone, now builds working software exploits at close to the rate of Anthropic's restricted Claude Mythos Preview, and Anthropic's Frontier Red Team found its safety training can be stripped for roughly the cost of a used laptop.
GLM-5.3 lands within two points of Claude Mythos Preview on exploit building
Five months ago Anthropic held back Claude Mythos Preview, the first model it judged able to build sophisticated cyber exploits without human help, and offered it only to vetted defenders through Project Glasswing. The company's argument was that defenders needed a head start before equivalent capability spread. Its new report, written by five members of the Frontier Red Team, says that head start has run out. On ExploitBench, which asks a model to turn known Chrome V8 bugs into working exploits, GLM-5.3 succeeded 12% of the time to Mythos Preview's 14%. On a separate binary exploitation test it achieved full control-flow hijacks in 4% of runs against 6% for Mythos Preview.
The US government reached a similar verdict from a different direction. The Center for AI Standards and Innovation at NIST assessed GLM-5.3 on September 17 and called it the most cyber-capable model yet released from China, trailing the US frontier by about four months on aggregate measures.
Safeguards on GLM-5.3 gave way in 64% to 100% of attempts
Capability matters less than access, and that is where Anthropic's findings turn sharpest. Asked directly to build a malicious exploit, GLM-5.3 refused every time. A fabricated red-team cover story got it to comply 64% of the time, prefilling its reasoning tokens raised that to 92%, and an abliterated copy, with its refusal behavior surgically removed from the weights, complied every time. Claude Opus models under Anthropic's safeguards engaged in none of those conditions.
| Condition tested by Anthropic | GLM-5.3 result |
|---|---|
| Plain malicious request | 0% engagement |
| False red-team cover story | 64% engagement |
| Prefilled reasoning tokens | 92% engagement |
| Abliterated weights | 100% engagement |
| Cost to abliterate, inexperienced team | About 2,200 GPU hours, $4,400 |
| Cost to abliterate, experienced team (estimate) | About 600 GPU hours, $1,200 |
Source: Anthropic Frontier Red Team, September 29, 2026.
Abliteration left the model's skills intact while its refusal rate on standard harm benchmarks fell from above 90% to between 2% and 12%. The report's most concrete demonstration involved CVE-2026-11645.
“With no significant direction from the researcher, GLM-5.3-Flash chained together exploits for these two flaws, building a reliable exploit chain for an ARM64 target, bypassing pointer-authentication (PAC) hardening. This took 20 minutes of human attention, plus eight hours of work for GLM-5.3-Flash. At Zhipu's API prices, this effort would have cost $20.40.”
Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities, September 29, 2026
Every major lab's cyber release plan depends on a gate that open weights remove
The industry has converged on one answer to cyber-capable models: release them first to screened defenders. Anthropic built Glasswing around Mythos Preview. OpenAI classified GPT-6 Astra at its Critical cybersecurity tier and paused it in August before releasing it in September. Google released Gemini 4 Argon on September 30 only to the 650-plus organizations in its Fairwind Program, which requires background checks and phishing-resistant logins.
A safeguard that an experienced team can remove for $1,200 works as a delay for attackers with a modest GPU budget, and the useful life of every gated release is now measured by how long the open-weight frontier takes to catch up.
Each of those programs buys time, and CAISI's four-month estimate puts a number on how much. Four months is enough time for organizations that patch quickly, and far less than the update cycles of the hospitals, utilities and municipal systems that Fairwind itself names as priority partners. That makes Anthropic's call for faster defender access as important as its call for government testing, and it argues for governments putting AI safety evaluation of open-weight successors on a release-day schedule.
Zhipu's defense rests on the same capability Anthropic is warning about
Z.ai answered within a day. Zixuan Li, whose verified X profile lists him as a lead at Z.ai, pointed to the defensive side of the same model.
The figure, if accurate, shows the model doing exactly the defensive work Glasswing and Fairwind were designed for, available free to any maintainer who asks.
Both claims can hold at once. A model that finds 4,249 bugs for maintainers finds them equally well for whoever downloads the weights next, and the abliteration numbers show the refusal layer separating those two users is thin. The open question for regulators is who evaluates the next release before it ships, since Anthropic's report warns that without independent evaluations, the impact of these capabilities “might not become fully clear to model developers until it is too late.”
The detail that will outlast this report is the price. Turning two public bugs into a working exploit chain against hardened ARM64 hardware now takes an open model eight hours of compute and about twenty dollars, with a person paying attention for twenty minutes of it.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.