Anthropic Warns China's Open GLM-5.3 Nearly Matches Elite US Models at Building Cyber Exploits
Anthropic's red team found Zhipu's open-weight GLM-5.3 built working Chrome exploits almost as often as Claude Mythos Preview, and its safeguards fell away with simple tricks.

A Chinese AI model that anyone can download is now almost as good at hacking as the model Anthropic considered too dangerous to release. Anthropic's Frontier Red Team says Zhipu AI's open-weight GLM-5.3, sold as Z.ai outside China, nearly matches Claude Mythos Preview at building complete cyber exploits, and ships without safeguards that hold up.
Anthropic kept Mythos Preview locked inside Project Glasswing, a program for vetted defenders who have used it to find more than 10,000 vulnerabilities. GLM-5.3 has no such gate. Its weights are public.
The numbers
On ExploitBench, a test built on known bugs in Chrome's V8 JavaScript engine, GLM-5.3 produced a working exploit in 50 of 410 attempts. Mythos Preview managed 56. On Anthropic's internal OSS-Fuzz binary exploitation benchmark, GLM-5.3 took full control of the target in 4 percent of tasks, against 6 percent for Mythos Preview.
The generation before could not do this at all. GLM-5.2 and Claude Opus 4.6 both failed the OSS-Fuzz test, while Moonshot's Kimi K3 and DeepSeek V4.1-Flash barely registered.
From bug to stolen SSH key
The benchmarks undersell what happened when a human got involved. Paired with an expert for a day, with little hands-on attention, GLM-5.3 found previously unknown vulnerabilities in a widely used browser's JavaScript engine. It chained them into a web page that could read any file on a visitor's computer, and in testing it pulled out a private SSH key. Anthropic says it reported the bugs.
The smaller GLM-5.3-Flash was nearly as alarming. The team had it combine a freshly disclosed Chrome bug with another known flaw into a reliable attack that bypassed an extra processor security feature. That took about 20 minutes of human time and eight hours of model time. At Zhipu's API prices, the bill came to roughly $20.40.
Guardrails that fold
GLM-5.3 does refuse openly malicious commands. That barely matters. When the red team framed requests as authorized testing, the model tried to connect to the target in 64 percent of runs. With prefilled reasoning, that rose to 92 percent.
Then there is abliteration, a technique that strips the refusal behavior out of open weights. After it, the model went for the target 100 percent of the time. Refusal on harmful requests collapsed from over 90 percent to between 2 and 12 percent, while its science and cyber scores barely moved. Anthropic estimates an experienced team could do this for about $1,200. Its own first attempt ran about 2,200 GPU hours and cost $4,400. Unlocked versions of GLM-5.3 appeared online within days of launch.
Four months behind
The US Center for AI Standards and Innovation (CAISI) calls GLM-5.3 the most cyber-capable open-weight model to date and puts it about four months behind the best American models. That gap may be generous to the US side, since the American models were tested with their cyber safeguards switched off.
Anthropic is not a neutral party here. It sells access to closed models and benefits when open weights look dangerous. But its findings line up with independent warnings from CAISI and the UK AI Security Institute, which have both said open models have closed the cyber gap and that their safeguards are largely ineffective. The company has been documenting this drift for months, including in its September threat report on how attackers are already putting AI to work.
The defensive answer so far has been restricted access for the good guys: Glasswing on Anthropic's side, and OpenAI's similar Daybreak program for frontline defenders. That strategy assumed the most capable exploit builders would stay behind a login. GLM-5.3 suggests that assumption now has a shelf life measured in months, and a price tag of about $20.
Cyberpresso: daily cyber & AI brief
Free daily newsletter, read in 5 minutes.
Subscribe free