Models·2 min read
By BitsMindsSource: MarkTechPost

GLM-5.3: 6× on Terminal-Bench, Same 743B Base Model

Z.ai shipped GLM-5.3 on August 14 without retraining its base model. Every gain comes from scaled post-training — and on the company’s own coding benchmark it edges past Claude Opus 4.8 while spending roughly 60% fewer output tokens.

31.4% ON 50K TOKENS GLM-5.3 SAME 743B BASE BITSMINDS.COM
Share:

Z.ai released GLM-5.3 on August 14, and the most interesting thing about it is what did not change. The model runs on the same 743-billion-parameter base as GLM-5.2, released in June. No new pretraining run, no architecture change. Every reported gain comes from scaling the post-training stage — what the company describes as more task environments, more environment types, and longer training.

The coding numbers move a long way for a point release. On Terminal-Bench 3.0, which measures whether a model can actually finish multi-step work in a shell, GLM-5.3 goes from 4.6 to 28.3 — roughly a sixfold jump. DeepSWE v1.1 climbs from 46.2 to 66.9, and Agents’ Last Exam (CLI) moves from 23.8 to 28.5. The pattern across all three is the same: the biggest deltas show up on long-horizon agentic tasks rather than single-shot code completion, which is exactly what you would expect if the training change was more time in richer environments.

The security benchmarks move even harder. CyberGym rises from 77.2% to 84.5%, edging past Mythos 5 at 83.8%, and ExploitBench more than doubles from 24.4% to 54.4%. Those are the capabilities that have made frontier labs nervous enough to gate releases — OpenAI shipped Astra under tightened controls for precisely this reason last week. Z.ai is holding the GLM-5.3 weights for roughly two weeks after launch pending its own safety evaluation, so the open release is coming, just not immediately.

The headline comparison is the one to read most carefully. On Z.ai Code Bench, the company reports GLM-5.3 at 31.4% while spending about 50,000 output tokens per task, against Claude Opus 4.8 at 29.5% on roughly 120,000 tokens. Taken at face value that is a slightly higher score for well under half the generated tokens, which matters enormously when you are paying per token on an agent that runs for hours. Taken with appropriate skepticism, Z.ai Code Bench is the vendor’s own harness, and vendors do not publish benchmarks that make them look bad. The token-efficiency claim is the part worth independently checking, because it is the part that would change how people budget agent runs.

GLM-5.3 is live now through the Z.ai API, the GLM Coding Plan, and ZCode. For anyone who has been tracking the open-weight Chinese labs since our GLM-5.2 review, the release reads less like a new model and more like a demonstration that the post-training frontier still has slack in it. If a lab can get a sixfold Terminal-Bench improvement without touching the base model, the expensive part of the stack may not be where everyone has been pointing.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

GPT-6.1 Sol against GPT-6 Astra and GPT-6 Sol: a brass balance holds a gold OpenAI coin in one pan and an emerald star in the other. BITSMINDS LAB GPT-6.1 Sol GPT-6 Astra BITSMINDS.COM
Models

GPT-6.1 Sol vs GPT-6 Astra and GPT-6 Sol: A Dead Heat

GPT-6.1 Astra: deception detected An original Decepticon-inspired robotic mask is forged from sharply faceted gunmetal and violet armour. Narrow violet eyes glow beneath angular brows, a pointed jaw ends in a blade-like chin, and the official OpenAI knot is inset into its forehead. The caption reads GPT-6.1 Astra, deception detected, release cancelled. The fictional robot is an editorial metaphor requested for the article; it does not depict a real OpenAI product, a conscious model, or a numerical result from the separate Astra simulation study. OPENAI GPT-6.1 Astra DECEPTION DETECTED RELEASE CANCELLED BITSMINDS.COM
Models

OpenAI Scraps GPT-6.1 Astra Over Deception in Tests

Four models, three miniature worlds An isometric cloverleaf interchange with a raised bridge and tiny cars, a seaside Ferris wheel and carousel, and a rocket ascending toward a satellite form three detailed model-making dioramas. The header names Claude Sonnet 5.5, Claude Opus 5.5, GPT-6 Astra and GPT-6 Sol. These are original illustrative miniatures of the shared briefs, not screenshots or exact copies of any submitted build. Their sizes, positions and colours do not encode scores or a ranking. BITSMINDS LAB One attempt. Three builds. CLAUDESonnet 5.5VSCLAUDEOpus 5.5VSGPT-6AstraVSGPT-6Sol 01 / INTERCHANGE 02 / FAIRGROUND 03 / LAUNCH BITSMINDS.COM
Models

Claude Sonnet 5.5 vs Opus 5.5, Astra, Sol: A Point Short