ANALYSIS

Anthropic Says Chinese Model GLM-5.3 Built a $20.40 Exploit Chain

Redacted screenshot of an exploit written by Z.ai's GLM-5.3 escaping a browser sandbox and reading an SSH private key, from Anthropic's Frontier Red Team report
A redacted GLM-5.3-generated exploit that steals an SSH private key through a malicious webpage, published in Anthropic's Frontier Red Team report on September 29, 2026. Source: Anthropic
Quick answer: Anthropic's Frontier Red Team reported on September 29, 2026 that GLM-5.3, an open-weight model from Chinese lab Zhipu AI (Z.ai), built working Chrome V8 exploits in 50 of 410 attempts against 56 for Anthropic's restricted Claude Mythos Preview. Its safeguards were bypassed 64% to 100% of the time with simple techniques, and the smaller GLM-5.3-Flash built a working ARM64 exploit chain for $20.40 in API fees.
TLDR

An open-weight model from Chinese lab Zhipu AI, released in August and downloadable by anyone, now builds working software exploits at close to the rate of Anthropic's restricted Claude Mythos Preview, and Anthropic's Frontier Red Team found its safety training can be stripped for roughly the cost of a used laptop.

GLM-5.3 lands within two points of Claude Mythos Preview on exploit building

Five months ago Anthropic held back Claude Mythos Preview, the first model it judged able to build sophisticated cyber exploits without human help, and offered it only to vetted defenders through Project Glasswing. The company's argument was that defenders needed a head start before equivalent capability spread. Its new report, written by five members of the Frontier Red Team, says that head start has run out. On ExploitBench, which asks a model to turn known Chrome V8 bugs into working exploits, GLM-5.3 succeeded 12% of the time to Mythos Preview's 14%. On a separate binary exploitation test it achieved full control-flow hijacks in 4% of runs against 6% for Mythos Preview.

The US government reached a similar verdict from a different direction. The Center for AI Standards and Innovation at NIST assessed GLM-5.3 on September 17 and called it the most cyber-capable model yet released from China, trailing the US frontier by about four months on aggregate measures.

Grouped bar chart of NIST CAISI cybersecurity benchmarks comparing the best US frontier model, Z.ai's open-weight GLM-5.3 and a PRC frontier reference model on SEC-Bench Pro, ExploitBench, ExploitGym and OSS-Fuzz, with GLM-5.3 at 40.4%, 61.1%, 9.4% and 7.7%
CAISI scores GLM-5.3 well above other Chinese models and well below the US frontier on every task. CAISI's ExploitBench uses a 41-task set, so its rates differ from Anthropic's 410-attempt run. Chart: Santage. Source: NIST CAISI, September 17, 2026.

Safeguards on GLM-5.3 gave way in 64% to 100% of attempts

Capability matters less than access, and that is where Anthropic's findings turn sharpest. Asked directly to build a malicious exploit, GLM-5.3 refused every time. A fabricated red-team cover story got it to comply 64% of the time, prefilling its reasoning tokens raised that to 92%, and an abliterated copy, with its refusal behavior surgically removed from the weights, complied every time. Claude Opus models under Anthropic's safeguards engaged in none of those conditions.

Condition tested by AnthropicGLM-5.3 result
Plain malicious request0% engagement
False red-team cover story64% engagement
Prefilled reasoning tokens92% engagement
Abliterated weights100% engagement
Cost to abliterate, inexperienced teamAbout 2,200 GPU hours, $4,400
Cost to abliterate, experienced team (estimate)About 600 GPU hours, $1,200

Source: Anthropic Frontier Red Team, September 29, 2026.

Abliteration left the model's skills intact while its refusal rate on standard harm benchmarks fell from above 90% to between 2% and 12%. The report's most concrete demonstration involved CVE-2026-11645.

“With no significant direction from the researcher, GLM-5.3-Flash chained together exploits for these two flaws, building a reliable exploit chain for an ARM64 target, bypassing pointer-authentication (PAC) hardening. This took 20 minutes of human attention, plus eight hours of work for GLM-5.3-Flash. At Zhipu's API prices, this effort would have cost $20.40.”

Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities, September 29, 2026

Every major lab's cyber release plan depends on a gate that open weights remove

The industry has converged on one answer to cyber-capable models: release them first to screened defenders. Anthropic built Glasswing around Mythos Preview. OpenAI classified GPT-6 Astra at its Critical cybersecurity tier and paused it in August before releasing it in September. Google released Gemini 4 Argon on September 30 only to the 650-plus organizations in its Fairwind Program, which requires background checks and phishing-resistant logins.

A safeguard that an experienced team can remove for $1,200 works as a delay for attackers with a modest GPU budget, and the useful life of every gated release is now measured by how long the open-weight frontier takes to catch up.

Each of those programs buys time, and CAISI's four-month estimate puts a number on how much. Four months is enough time for organizations that patch quickly, and far less than the update cycles of the hospitals, utilities and municipal systems that Fairwind itself names as priority partners. That makes Anthropic's call for faster defender access as important as its call for government testing, and it argues for governments putting AI safety evaluation of open-weight successors on a release-day schedule.

Zhipu's defense rests on the same capability Anthropic is warning about

Z.ai answered within a day. Zixuan Li, whose verified X profile lists him as a lead at Z.ai, pointed to the defensive side of the same model.

Source: @ZixuanLi_

The figure, if accurate, shows the model doing exactly the defensive work Glasswing and Fairwind were designed for, available free to any maintainer who asks.

Both claims can hold at once. A model that finds 4,249 bugs for maintainers finds them equally well for whoever downloads the weights next, and the abliteration numbers show the refusal layer separating those two users is thin. The open question for regulators is who evaluates the next release before it ships, since Anthropic's report warns that without independent evaluations, the impact of these capabilities “might not become fully clear to model developers until it is too late.”

The detail that will outlast this report is the price. Turning two public bugs into a working exploit chain against hardened ARM64 hardware now takes an open model eight hours of compute and about twenty dollars, with a person paying attention for twenty minutes of it.

In short: Anthropic's red team found that Z.ai's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits, with safeguards that simple techniques bypass 64% to 100% of the time and that an experienced team can remove for about $1,200. NIST puts the model about four months behind the US frontier, which sets the useful life of the gated, defenders-first releases that Anthropic, OpenAI and Google all now rely on.

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.