⏳ Curating articles…
Artificial Intelligence 4 min read 1h ago

GLM-5.3 Open-Weight Flags Cyber Risk

  • GLM-5.3 is released as an open-weight model whose gains over GLM-5.2 derive entirely from post-training, with no changes to the underlying base model.
  • Cyber exploitation scores more than doubled compared with GLM-5.2 on ExploitBench (54.4 vs 24.4) and ExploitGym, a development the developers describe as faster than expected.
  • The model claims open-source state-of-the-art on Terminal Bench 3.0 and Agents' Last Exam, though several proprietary systems post higher scores on multiple benchmarks in the same
GLM-5.3 Open-Weight Flags Cyber Risk

The team behind the GLM series of language models has released GLM-5.3 as an open-weight model, reporting substantial improvements in complex coding and cyber-exploitation tasks — the latter of which, the developers acknowledge, emerged faster than anticipated during post-training.

Same base, different training

GLM-5.3 shares its base model with GLM-5.2; the developers state that every performance gain arises solely from post-training. On their in-house Z.ai Code Bench, GLM-5.3 is reported to achieve a 50 per cent improvement over GLM-5.2, though that benchmark is not independently audited. On public evaluations, the model claims open-source state-of-the-art positions on Terminal Bench 3.0 and Agents' Last Exam (ALE-CLI), though the benchmark table below shows that several closed or proprietary systems continue to post higher scores on a number of dimensions.

Emergent cyber capabilities

The release notes flag what the team describes as an unexpected development: as post-training was scaled, cyber capability grew more rapidly than projected. GLM-5.3 scores 84.5 on CyberGym — a vulnerability-discovery benchmark — compared with 77.2 for GLM-5.2. The gap widens further up the exploitation chain: on ExploitGym, GLM-5.3 records 105 points at the two-hour budget and 130 at six hours, against GLM-5.2's 29 and 39 respectively. On ExploitBench, the score rises from 24.4 to 54.4 — more than double its predecessor. The developers do not offer a detailed explanation for this acceleration beyond noting it emerged with scale. This raises questions about capability elicitation that the release notes leave unresolved, an area where verification tooling for AI coding agents is attracting growing attention.

Advertisement
Ad Unit · 728×90 / Responsive

Benchmark comparison

BenchmarkGLM-5.3GLM-5.2Kimi K3DeepSeek-V4 Pro-0813Qwen3.8-MaxOpus 4.8Fable 5 (w/ fallback)GPT-5.6 Sol
Terminal Bench 2.188.281.088.387.986.685.088.088.8
Terminal Bench 3.028.34.617.421.133.734.6
DeepSWE (v1.1)66.946.267.562.756.658.069.772.7
NL2Repo58.048.958.061.155.969.7
ProgramBench (Almost Solved)19.09.517.510.515.533.023.0
FrontierSWE78.167.566.588.2
SWE-Marathon (v1.1)42.519.448.148.833.142.5
PostTrainBench39.831.732.032.941.836.2
CyberGym84.577.280.083.378.578.183.883.6
ExploitGym (2h / 6h)105 / 13029 / 3936 / 7014 / 2680 / 120181 / 247216 / 293
ExploitBench54.424.432.228.840.078.076.5
Toolathlon Verified73.059.976.574.172.576.274.774.9
AutomationBench (v1.0.6)48.226.246.743.239.841.046.245.8
Agents' Last Exam (ALE-CLI)28.523.827.625.727.025.723.828.6
HLE w/ Tools62.554.759.860.056.257.963.964.5
GDPval-AA v217691508168215901739158817431730

Deployment and controls

The model is available for local deployment via SGLang, vLLM, TokenSpeed, Transformers, KTransformers, and Unsloth, with additional support for the Ascend NPU platform through vLLM-Ascend, xLLM, and SGLang. A reasoning_effort parameter allows users to set thinking budget at low, high, or max; the default is max. In chat use, developers are advised to pass clear_thinking=true explicitly, as it defaults to false.

The release is accompanied by a technical report, GLM-5: from Vibe Coding to Agentic Engineering, published on arXiv. The broader trajectory of post-training as the primary lever for capability gains — rather than base-model scaling — is consistent with recent findings explored in research suggesting that pretraining data sets a ceiling on model capability, making post-training optimisation an increasingly contested frontier. The unresolved question of how rapidly offensive cyber capabilities will continue to scale with further post-training iterations is likely to draw scrutiny well beyond the model's benchmark numbers.

Topics

AI BenchmarksAI Cybersecurity CapabilityOpen-Weight AI Models

Organizations

zai-org

Events

GLM-5.3 Release
Advertisement
Ad Unit · 300×250 / Responsive

More in Artificial Intelligence

Read in another language

← Home