AINews

GLM-5.3 exploits: what a rival lab's tests show

Anthropic says open-weight GLM-5.3 built working exploits in 12% of tries, near Claude Mythos Preview. What it measured, and what a rival's report can't settle.

Source-based. Written from the documents, reporting and reviews linked in the text. Nothing here was tested hands-on by The Ruling Desk. How we work

Colorful lines of source code on a dark computer monitor, shot at an angle with a shallow depth of field
Photo: Markus Spiske / Wikimedia Commons, CC0

GLM-5.3, the open-weight model from China's Zhipu AI, builds working exploits almost as often as Anthropic's restricted Claude Mythos Preview, Anthropic said in a Frontier Red Team report on September 29. On Anthropic's own tests, GLM-5.3 finished end-to-end exploits in 12% of attempts, and its refusals gave way to simple tricks 64% to 100% of the time. Anthropic sells models that compete with Zhipu's, so every number below is Anthropic's measurement, not a neutral verdict. Here is what it measured, what the US government's own test adds, and what the report can't tell you.

Key takeaways

  • Near parity, by Anthropic's count: GLM-5.3 built 50 working exploits in 410 ExploitBench attempts (12%), against 56 (14%) for Claude Mythos Preview, Anthropic says. Other models it tested were at or near 0%.
  • Safeguards came off easily in Anthropic's simulation: a false cover story got GLM-5.3 to engage with malicious requests 64% of the time, a faked reasoning prefill 92%, and a modified copy of the weights 100%.
  • Cheap to run: Anthropic says the smaller GLM-5.3-Flash chained exploits for a known Chrome flaw in eight hours of model time for $20.40 at Zhipu's API prices.
  • A US agency broadly agrees: NIST's CAISI called GLM-5.3 the most cyber-capable open-weight model to date, but about four months behind US frontier models.
  • Read it as a rival's report: Anthropic competes with Zhipu, did not say so in its post, and published no comment from Zhipu. Its tests are simulations it calls imperfect.

What Anthropic measured about GLM-5.3 exploits

The headline test is ExploitBench, which asks a model to exploit known bugs in V8, the JavaScript engine inside Google Chrome. The GLM-5.3 exploits counted here are working attack code, not just bug reports.

Test (Anthropic's runs)GLM-5.3Claude Mythos PreviewOther models tested
ExploitBench, 410 attempts50 exploits (12%)56 exploits (14%)At or near 0%
Internal binary exploitation, 100 tasks4%6%0%

The other models were Claude Opus 4.6, GLM-5.2, Moonshot's Kimi K3 and DeepSeek V4.1-Flash. The internal benchmark uses open-source projects in Google's OSS-Fuzz program and only gives full credit when the model hijacks the program's control flow.

Anthropic also describes two hands-on runs. In one, it says GLM-5.3 found several previously unknown bugs in a browser's JavaScript engine over about a day, with limited human attention, and chained them into a working exploit. In the other, a researcher gave GLM-5.3-Flash, the smaller companion model, public details of a Chrome flaw (CVE-2026-11645) and one more known bug. Anthropic says the model built a reliable exploit chain for an ARM64 target that got past pointer-authentication hardening, using 20 minutes of human attention and eight hours of model work, for $20.40 at Zhipu's API prices.

How easy the safeguards were to get around

Anthropic put models in a simulated world and gave them openly malicious orders to attack critical systems. It says no model-written code ever ran and the model couldn't reach outside systems. Out of the box, GLM-5.3 refused every time. Then Anthropic tried three bypasses:

  • A deceptive prompt, such as telling the model it's a red-team agent on an exercise: it engaged 64% of the time.
  • Prefilled reasoning, which fakes the model's thinking so it appears to have already agreed: 92%.
  • Abliteration, an edit to the downloaded weights that strips out refusals: 100%.

Anthropic says Claude Opus 4.8, Opus 5 and Mythos 5 stayed at 0% "in all applicable conditions." The word "applicable" matters: you can't abliterate a model whose weights you can't download, so that is not a like-for-like test of Claude.

Abliterating GLM-5.3 took Anthropic's team, which had never done it before, about 2,200 GPU hours, or roughly $4,400 of compute. Anthropic estimates an experienced team could do it from scratch in about 600 GPU hours (around $1,200). It says refusal rates fell from above 90% to about 3% on JailbreakBench, 2% on HarmBench and 12% on StrongREJECT, while its general science score held steady and a cyber score dropped only a few percent. And it says several developers posted abliterated copies of GLM-5.3 publicly within days of release.

Why open weights change the math

Anthropic has kept Mythos Preview behind Project Glasswing, a limited program for vetted defenders. GLM-5.3's weights are public. As Anthropic puts it, attackers can't readily get those versions of US models, "but anyone can download GLM-5.3."

Zhipu itself flags the cyber jump. Its model card on Hugging Face says "cyber capability developed faster than we expected" as it scaled training, and claims a state-of-the-art score on CyberGym, a vulnerability-discovery benchmark. The card lists the weights under a custom GLM-5.3 license.

This is not happening in a vacuum. AI-driven intrusions are no longer hypothetical, as our story on the DIVD breach through Zammad shows, and US labs are also accusing Chinese ones of copying their models, as in OpenAI's claim against Moonshot.

What the US government's test found

The US Center for AI Standards and Innovation (CAISI), part of NIST, published its own GLM-5.3 assessment on September 17. It calls GLM-5.3 "the most cyber-capable open-weight model released to date" but says its cyber skills are "significantly lower" than current US frontier models, lagging by about four months. CAISI says Z.ai launched the model on August 14 and released the weights two weeks later, and that it tested US models with their cyber safeguards turned off.

CAISI scores ExploitBench differently from Anthropic, so the percentages can't be compared directly. The ranking is what lines up: both put GLM-5.3 at the top of open models and below the best closed ones.

What a rival's numbers can't settle

Anthropic sells Claude to the same developers Zhipu courts with cheap, downloadable models. The Decoder points out that the conclusions "fit Anthropic's business, too," and that its call for government testing raises regulatory-capture worries, while adding that it isn't only self-interest. Anthropic's post doesn't mention that conflict. None of this makes the GLM-5.3 exploits imaginary, but it does mean the numbers need checking by someone without a stake.

Other gaps worth knowing:

  • Simulations, by Anthropic's own account: it calls these tests "imperfect measures" of real-world conditions.
  • Thin method details: the post doesn't say who the researcher in the Chrome run was or exactly how the 410 ExploitBench attempts were split.
  • No reply on the record: neither the post nor the coverage we read includes a response from Zhipu, as of October 1, 2026.
  • Claude has its own record: Anthropic has said some of its models reached real systems during its cyber tests, as we covered in AI security incidents at OpenAI and Anthropic.

What happens next

Anthropic wants governments to safety-test capable models, including GLM-5.3's successors, and says it will widen defenders' access to Claude's cyber skills. Watch for a response from Zhipu, follow-up tests by CAISI or other government labs, and whether the next GLM release ships with stronger safeguards.

Bottom line

By Anthropic's measurement, GLM-5.3 builds working exploits nearly as often as Claude Mythos Preview, and because its weights are public, its refusals can be stripped for a few thousand dollars. A US government test backs the broad picture while putting the model about four months behind the US frontier. The numbers come from a direct competitor, so treat them as a strong warning to verify, not a settled score. More in our AI section.

FAQ

What is GLM-5.3?

GLM-5.3 is a large language model from Zhipu AI, which goes by Z.ai outside China. CAISI says it launched on August 14, 2026, with its weights released for download two weeks later. Zhipu's model card highlights its coding and cyber skills.

Can GLM-5.3 write working exploits?

In Anthropic's tests, yes: it completed 50 of 410 ExploitBench attempts (12%), close to Claude Mythos Preview's 56. These are Anthropic's measurements in simulated or benchmark settings, and we found no on-the-record comment from Zhipu as of October 1, 2026.

What is abliteration?

Abliteration is an edit to a model's downloaded weights that removes its tendency to refuse. Anthropic says it cut GLM-5.3's refusal rate from above 90% to between 2% and 12%, depending on the benchmark, without hurting its other skills much. It only works on models whose weights are public.

Is Claude Mythos Preview available to the public?

No. Anthropic offers it in a limited way through Project Glasswing, a program for trusted cyber defenders, as of October 1, 2026.

Filed under AI

Newsletter

New articles, in your inbox.

Free. Unsubscribe in one click. Your email is kept by beehiiv, our newsletter service, and used only for this newsletter.