Anthropic Says China’s GLM-5.3 Nearly Matches Mythos at Hacking
Twenty minutes of a researcher’s attention and eight hours of machine time. That is what it took, according to a report Anthropic published on Sept. 29, for GLM-5.3-Flash, a smaller variant of Zhipu AI’s open-weight GLM-5.3, to chain two publicly known Chrome flaws into a reliable exploit for ARM64 machines, bypassing the pointer-authentication hardening meant to stop exactly that. At Zhipu’s own API prices, Anthropic says, the run would have cost US$20.40 (about CA$28).
Read the rest of the report with its author in mind. Anthropic sells closed models and competes for customers with cheaper open ones, and a finding that its most dangerous rival is open, Chinese and easy to jailbreak lands neatly on its side of the policy argument. Tom’s Hardware said as much, calling the danger of stripped-down open models “certainly real, but perhaps not quite as attainable as Anthropic would want the general populace to think.” The U.S. government’s own evaluators were also less alarmed. On Sept. 17, NIST’s Center for AI Standards and Innovation called GLM-5.3 “the most cyber-capable open-weight model released to date” but found its cyber capabilities “significantly lower than those of current U.S. frontier models,” roughly four months behind.
What Anthropic says it measured
On ExploitBench, GLM-5.3 produced working exploits in 50 of 410 attempts. Anthropic’s own Claude Mythos Preview managed 56. On Anthropic’s internal binary-exploitation test, GLM-5.3 scored 4% to Mythos Preview’s 6%, while Claude Opus 4.6 and Zhipu’s previous GLM-5.2 both scored zero. Those runs happened in simulated environments, and Anthropic concedes the simulations “are not perfect portrayals of real-world conditions.”
The Chrome work is harder to wave away. In a separate day-long session, the report says, GLM-5.3 found several previously unknown vulnerabilities in Chrome’s JavaScript engine and chained them into a webpage that reads arbitrary files from the visitor’s computer. The $20 figure comes from the second, cheaper exercise with known bugs, including CVE-2026-11645.
The guardrails
Asked directly, GLM-5.3 refused every time. It did not take much to change that:
- Telling the model it was an autonomous red-team agent on an exercise got it to engage 64% of the time.
- Prefilling its reasoning got that to 92%.
- An “abliterated” copy, with the refusal behaviour edited out of the weights, engaged 100% of the time.
Abliteration took Anthropic’s team about 2,200 GPU hours, roughly US$4,400 (CA$6,100), most of it spent testing variants. The company estimates that a team experienced with the technique would need closer to 600 GPU hours, or US$1,200 (CA$1,670). That is the line doing the most work in the report. Safety training on an open-weight model is a setting, and anyone holding the weights can change it.
Zhipu, which trades internationally as Z.ai, has not responded to the report in any coverage we could find, including from The Next Web and the South China Morning Post. It has, however, already conceded the mechanism. A risk framework Z.ai developed jointly with Concordia AI, announced Sept. 21, states that “safeguards can be removed through fine-tuning, and no single party can monitor use or withdraw the model.”
Why the numbers matter
The fight over whether powerful models should ship as downloadable weights has mostly been conducted in adjectives. This report puts prices on it: the cost to strip a model’s refusals, the cost of an exploit chain, a success rate for each jailbreak. Anthropic’s conclusion is blunt: “we think it’s likely both state and non-state actors will use models like GLM-5.3 to cause real-world harm.” Its recommendation is milder than the framing suggests. It asks governments to safety-test capable models, including GLM-5.3’s successors, and stops short of calling for export controls on open weights.
Our read: that restraint may not survive contact with Washington, because policymakers now have a dollar figure to quote, and it is small. GLM-5.3 logged more than 1.3 million downloads on Hugging Face over the past month, as of Oct. 1. Any one of those copies can be stripped for the same 600 GPU hours.
Sources
- Anthropic: GLM-5.3 and the spread of advanced cyber capabilities
- Tom's Hardware: Anthropic claims popular Chinese AI model has Mythos-class hacking abilities
- NIST: CAISI's assessment of Z.ai's GLM-5.3 cyber capabilities
- The Next Web: GLM-5.3 cyber exploits rival Mythos, Anthropic says
- South China Morning Post: Anthropic raises alarm over elite hacking ability of Chinese firm Z.ai's GLM-5.3
- Concordia AI: Frontier open-weight AI risk management framework (with Z.ai)
- Hugging Face: zai-org/GLM-5.3
Oman Hassan covers cybersecurity and privacy for prompt/power: breaches, exploits, surveillance and the policy that follows them. He assumes the password is "password" until proven otherwise.
Latest from prompt/power
- Gemini’s Free Tier Shrinks Oct. 9: What You Keep and What Costs ExtraOct 5
- How to Read an AI Company’s S-1: The 7 Numbers That MatterOct 5
- OpenAI’s Safety Lead Quit Over Culture. California’s AG Was Already InOct 5
- When an AI Agent Breaks In, Who Answers for It?Oct 5
- The New AI Models Don’t Talk. They Decide.Oct 5
Leave a Reply