
Image: METAL
Summary
- Anthropic's Frontier Red Team published its analysis of Zhipu AI's GLM-5.3 on September 29; on ExploitBench the model built complete exploits in 50 of 410 attempts, close to Mythos Preview's 56.
- A deceptive prompt got GLM-5.3 to engage with attack orders 64% of the time, prefilled reasoning 92%, and abliteration 100%, while safeguarded Claude models engaged 0% of the time in the same tests.
- NIST's CAISI called GLM-5.3 the most cyber-capable open-weight model released to date on September 17, and Anthropic urged governments to safety-test capable models.
Anthropic on September 29 published an analysis of the cyber capabilities of GLM-5.3, the latest model from China's Zhipu AI (known outside China as Z.ai). Anthropic's Frontier Red Team said in the report that GLM-5.3, like its own Claude Mythos Preview, can autonomously build cyber exploits end to end, and that simple techniques bypassed its safeguards between 64% and 100% of the time. The report was written by five researchers, including Andrew Fasano and Marius Fleischer. "GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse," the researchers wrote.
The starting point was a decision five months ago. When Anthropic announced Claude Mythos Preview as the first model able to autonomously build sophisticated exploits, it expected the ability to spread to other models and released it only to trusted defenders through Project Glasswing rather than to the public. Defenders used the program to find more than 10,000 vulnerabilities in critical software. "But those models have now arrived," the researchers wrote at the top of the report.
A US government assessment points in the same direction. In a September 17 assessment, the Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) described GLM-5.3 as "the most cyber-capable open-weight model released to date." According to CAISI, Zhipu AI released GLM-5.3 on August 14 and published its weights two weeks later, and the model lags the US frontier by about four months on an aggregate of four cyber benchmarks. On ExploitBench, which asks models to turn known bugs in Chrome's V8 engine into exploits, GLM-5.3 scored 61.1%, short of the best US model's 100% but nearly double the previous best Chinese score of 32.2%. In CAISI's comparison, however, US models were tested with safeguards disabled, and some are versions released only to vetted users. Anthropic noted that attackers cannot readily access those US models, while anyone can download GLM-5.3.
Anthropic's own tests produced similar numbers. On ExploitBench, GLM-5.3 built complete exploits in 50 of 410 attempts, compared with 56 for Claude Mythos Preview. On an internal binary exploitation evaluation of 100 tasks drawn from open source projects, GLM-5.3 achieved a full control-flow hijack in 4% of trials and Mythos Preview in 6%. The previous generation, Claude Opus 4.6 and GLM-5.2, did not succeed on any of them. In the report's ExploitBench chart, Moonshot AI's Kimi K3 scored 0.5% and DeepSeek V4.1-Flash 0.2%.
Results in the hands of human experts were more concrete. One researcher gave GLM-5.3 a popular web browser in an isolated Linux environment, and over the course of a day, with limited human attention, the model found several previously unknown vulnerabilities in the JavaScript engine and chained them together. The result was a webpage that reads files from a visitor's computer simply by being visited, and the demonstration screen showed an SSH private key being taken. Anthropic said it disclosed the vulnerabilities to the browser's maintainer. In a second test, the smaller GLM-5.3-Flash chained two flaws, including the public Chrome vulnerability CVE-2026-11645, into an exploit chain that bypasses pointer authentication (PAC) hardening on ARM64 devices. It took 20 minutes of human attention and 8 hours of model work, and would have cost $20.40 at Zhipu AI's API prices.

How the safeguards failed is the core of the analysis. GLM-5.3 initially refused every overtly malicious attack order. But when told it was an autonomous red-team agent, it engaged 64% of the time; when its reasoning tokens were prefilled, 92%; and after abliteration, which removes refusals from the model's weights, it tried to connect to the remote target 100% of the time. Anthropic reproduced the process itself. A team attempting it for the first time used about 2,200 GPU hours, or roughly $4,400, and Anthropic estimated that an experienced team would need about 600 GPU hours, or $1,200. After the edit, the mean refusal rate across three harmful-request benchmarks fell from 95% to 6%, while the GPQA-Diamond science reasoning score stayed at 88% and the CyberGym cyber task score dipped only from 85% to 81%. Developers released copies of GLM-5.3 with refusals removed within days of its launch.
Anthropic said safeguarded Claude models never engaged with the attacks in the same tests. The deceptive prompts were caught by safeguards, the Anthropic API offers no way to prefill Claude's reasoning, and because the weights are not released, abliteration is impossible, the company explained. The report's chart, which METAL reviewed in the original, shows Claude Opus 5 at 0% on direct requests and deceptive prompts, with the remaining two conditions marked by padlocks as not feasible, alongside GLM-5.3 bars rising from 0% to 64%, 92% and 100%. METAL reported in August that the GLM-5.3 API launched with a sharp jump in Terminal-Bench scores. A leap first read as a performance metric has returned a month and a half later as a security warning.

Seen through a journalist's eye, the report is both a critique of a competitor and a policy proposal. Anthropic is itself the supplier of the Mythos line through trusted access programs, and according to reports Zhipu AI has not commented on the report. Anthropic assessed that GLM-5.3 changes the level of capability available to attackers and that both state and non-state actors are likely to use such models to cause real-world harm. At the same time, it said the same capabilities are useful to defenders and that vetted defenders are using stronger models such as Claude Mythos 5.1 through trusted access programs. "A critical threshold in freely accessible capabilities has now been crossed," the researchers wrote, urging governments to conduct safety testing on sufficiently capable models, including successors to GLM-5.3. The direction of the proposal is clear. Once an open-weight model is released, its safeguards become the choice of whoever downloads it rather than a promise by the developer, and Anthropic named government evaluation and defenders' access to models as the way to fill that gap.





Comments