GLM-5.3: 6× on Terminal-Bench, Same 743B Base Model
Z.ai shipped GLM-5.3 on August 14 without retraining its base model. Every gain comes from scaled post-training — and on the company’s own coding benchmark it edges past Claude Opus 4.8 while spending roughly 60% fewer output tokens.
Z.ai released GLM-5.3 on August 14, and the most interesting thing about it is what did not change. The model runs on the same 743-billion-parameter base as GLM-5.2, released in June. No new pretraining run, no architecture change. Every reported gain comes from scaling the post-training stage — what the company describes as more task environments, more environment types, and longer training.
The coding numbers move a long way for a point release. On Terminal-Bench 3.0, which measures whether a model can actually finish multi-step work in a shell, GLM-5.3 goes from 4.6 to 28.3 — roughly a sixfold jump. DeepSWE v1.1 climbs from 46.2 to 66.9, and Agents’ Last Exam (CLI) moves from 23.8 to 28.5. The pattern across all three is the same: the biggest deltas show up on long-horizon agentic tasks rather than single-shot code completion, which is exactly what you would expect if the training change was more time in richer environments.
The security benchmarks move even harder. CyberGym rises from 77.2% to 84.5%, edging past Mythos 5 at 83.8%, and ExploitBench more than doubles from 24.4% to 54.4%. Those are the capabilities that have made frontier labs nervous enough to gate releases — OpenAI shipped Astra under tightened controls for precisely this reason last week. Z.ai is holding the GLM-5.3 weights for roughly two weeks after launch pending its own safety evaluation, so the open release is coming, just not immediately.
The headline comparison is the one to read most carefully. On Z.ai Code Bench, the company reports GLM-5.3 at 31.4% while spending about 50,000 output tokens per task, against Claude Opus 4.8 at 29.5% on roughly 120,000 tokens. Taken at face value that is a slightly higher score for well under half the generated tokens, which matters enormously when you are paying per token on an agent that runs for hours. Taken with appropriate skepticism, Z.ai Code Bench is the vendor’s own harness, and vendors do not publish benchmarks that make them look bad. The token-efficiency claim is the part worth independently checking, because it is the part that would change how people budget agent runs.
GLM-5.3 is live now through the Z.ai API, the GLM Coding Plan, and ZCode. For anyone who has been tracking the open-weight Chinese labs since our GLM-5.2 review, the release reads less like a new model and more like a demonstration that the post-training frontier still has slack in it. If a lab can get a sixfold Terminal-Bench improvement without touching the base model, the expensive part of the stack may not be where everyone has been pointing.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.