GLM-5.3 Open-Weight Flags Cyber Risk
- GLM-5.3 is released as an open-weight model whose gains over GLM-5.2 derive entirely from post-training, with no changes to the underlying base model.
- Cyber exploitation scores more than doubled compared with GLM-5.2 on ExploitBench (54.4 vs 24.4) and ExploitGym, a development the developers describe as faster than expected.
- The model claims open-source state-of-the-art on Terminal Bench 3.0 and Agents' Last Exam, though several proprietary systems post higher scores on multiple benchmarks in the same
The team behind the GLM series of language models has released GLM-5.3 as an open-weight model, reporting substantial improvements in complex coding and cyber-exploitation tasks — the latter of which, the developers acknowledge, emerged faster than anticipated during post-training.
Same base, different training
GLM-5.3 shares its base model with GLM-5.2; the developers state that every performance gain arises solely from post-training. On their in-house Z.ai Code Bench, GLM-5.3 is reported to achieve a 50 per cent improvement over GLM-5.2, though that benchmark is not independently audited. On public evaluations, the model claims open-source state-of-the-art positions on Terminal Bench 3.0 and Agents' Last Exam (ALE-CLI), though the benchmark table below shows that several closed or proprietary systems continue to post higher scores on a number of dimensions.
Emergent cyber capabilities
The release notes flag what the team describes as an unexpected development: as post-training was scaled, cyber capability grew more rapidly than projected. GLM-5.3 scores 84.5 on CyberGym — a vulnerability-discovery benchmark — compared with 77.2 for GLM-5.2. The gap widens further up the exploitation chain: on ExploitGym, GLM-5.3 records 105 points at the two-hour budget and 130 at six hours, against GLM-5.2's 29 and 39 respectively. On ExploitBench, the score rises from 24.4 to 54.4 — more than double its predecessor. The developers do not offer a detailed explanation for this acceleration beyond noting it emerged with scale. This raises questions about capability elicitation that the release notes leave unresolved, an area where verification tooling for AI coding agents is attracting growing attention.
Benchmark comparison
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | DeepSeek-V4 Pro-0813 | Qwen3.8-Max | Opus 4.8 | Fable 5 (w/ fallback) | GPT-5.6 Sol |
|---|---|---|---|---|---|---|---|---|
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 87.9 | 86.6 | 85.0 | 88.0 | 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | – | – | 21.1 | 33.7 | 34.6 |
| DeepSWE (v1.1) | 66.9 | 46.2 | 67.5 | 62.7 | 56.6 | 58.0 | 69.7 | 72.7 |
| NL2Repo | 58.0 | 48.9 | 58.0 | 61.1 | 55.9 | 69.7 | – | – |
| ProgramBench (Almost Solved) | 19.0 | 9.5 | 17.5 | – | 10.5 | 15.5 | 33.0 | 23.0 |
| FrontierSWE | 78.1 | 67.5 | – | – | – | 66.5 | 88.2 | – |
| SWE-Marathon (v1.1) | 42.5 | 19.4 | 48.1 | – | – | 48.8 | 33.1 | 42.5 |
| PostTrainBench | 39.8 | 31.7 | 32.0 | – | – | 32.9 | 41.8 | 36.2 |
| CyberGym | 84.5 | 77.2 | 80.0 | 83.3 | 78.5 | 78.1 | 83.8 | 83.6 |
| ExploitGym (2h / 6h) | 105 / 130 | 29 / 39 | 36 / 70 | – | 14 / 26 | 80 / 120 | 181 / 247 | 216 / 293 |
| ExploitBench | 54.4 | 24.4 | 32.2 | – | 28.8 | 40.0 | 78.0 | 76.5 |
| Toolathlon Verified | 73.0 | 59.9 | 76.5 | 74.1 | 72.5 | 76.2 | 74.7 | 74.9 |
| AutomationBench (v1.0.6) | 48.2 | 26.2 | 46.7 | 43.2 | 39.8 | 41.0 | 46.2 | 45.8 |
| Agents' Last Exam (ALE-CLI) | 28.5 | 23.8 | 27.6 | 25.7 | 27.0 | 25.7 | 23.8 | 28.6 |
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 60.0 | 56.2 | 57.9 | 63.9 | 64.5 |
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1590 | 1739 | 1588 | 1743 | 1730 |
Deployment and controls
The model is available for local deployment via SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth, with additional support for the Ascend NPU platform through vLLM-Ascend, xLLM, and SGLang. A reasoning_effort parameter allows users to set thinking budget at low, high, or max; the default is max. In chat use, developers are advised to pass clear_thinking=true explicitly, as it defaults to false.
The release is accompanied by a technical report, GLM-5: from Vibe Coding to Agentic Engineering, published on arXiv. The broader trajectory of post-training as the primary lever for capability gains — rather than base-model scaling — is consistent with recent findings explored in research suggesting that pretraining data sets a ceiling on model capability, making post-training optimisation an increasingly contested frontier. The unresolved question of how rapidly offensive cyber capabilities will continue to scale with further post-training iterations is likely to draw scrutiny well beyond the model's benchmark numbers.
Topics
Organizations
Events
More in Artificial Intelligence
→
Artificial Intelligence
AI Agents Turn Bug Rumours Into Exploits
Artificial Intelligence
Racter: AI Prose Generator From 1984
Artificial Intelligence
AI Agents Claim Five Maths Firsts
Artificial Intelligence
Conduct: Open-Source AI Agent Guardrails
Artificial Intelligence
AI Assists World-First Brain Surgery
Artificial Intelligence
Google Gemini Scrubs Ums from Voice Text