Zhipu GLM-5.3 Review 2026: Post-Training Only, Open-Source Coding #1, AA 60, CyberGym #1, $1.40/M
Verdict first: GLM-5.3 is the strongest argument yet that post-training scaling beats scaling parameters. Zhipu kept the exact same 743B MoE base as GLM-5.2 and still lifted coding by ~50% — Terminal-Bench 3.0 jumps from 4.6 to 28.3 (open-source #1), DeepSWE hits 66.9 (past DeepSeek V4 Pro), and an unexpected cybersecurity ability emerges with CyberGym at 84.5% (global #1). Its Artificial Analysis Intelligence Index of 60 ties Kimi K3 for the top open-source score, at the lowest single-task cost among frontier flagships. If you code or run agents and want open weights near the frontier, this is the one to watch.
GLM-5.3 — Zhipu's open-source coding flagship, powered purely by post-training
What Is GLM-5.3?
Released on August 14, 2026 with its API going live August 19, GLM-5.3 is Zhipu AI's (Z.ai) latest flagship — and its most provocative. The architecture and parameter count are identical to GLM-5.2 (743B MoE, ~40B active). The entire leap comes from post-training scaling, a move Zhipu summarizes bluntly: "Scaling post-training is all we did."
The result: open-source #1 coding, a frontier-band Intelligence Index of 60, and an unexpected emergence of cybersecurity capability that has security teams paying attention.
Specifications at a Glance
| Item | Value |
|---|---|
| Vendor | Zhipu AI (Z.ai) |
| Release / API live | August 14 / August 19, 2026 |
| Positioning | Open-source coding flagship, post-training only |
| Architecture | Sparse Mixture-of-Experts |
| Total / active params | 743B / ~40B active |
| Context window | 1M input / 128K output |
| Inputs | Text only |
| API price | $1.40 / $4.40 per M tokens (unchanged from 5.2) |
| Open weights | ~Aug 28, MIT license |
| Throughput | ~145–168 tokens/sec |
Pricing: Cheapest Frontier Cost
| Item | GLM-5.3 | Note |
|---|---|---|
| Input | $1.40 / M tokens | Unchanged from GLM-5.2 |
| Output | $4.40 / M tokens | — |
| Single-task cost | Lowest among frontier flagships | Artificial Analysis |
While DeepSeek and others are raising prices, GLM-5.3 holds its rate. And its token efficiency is the kicker: in High mode, GLM-5.3 scored 31.4% using ~50K tokens, while Claude Opus 4.8 needed ~120K tokens for 29.5% — half the tokens, higher completion. For heavy users, that compounds into dramatically lower bills.
Benchmarks: Post-Training Explosion
| Benchmark | 5.2 | 5.3 | Highlight |
|---|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3 | 6x, open-source #1 |
| DeepSWE v1.1 | 46.2 | 66.9 | Beats DeepSeek V4 Pro (62.7) |
| Agents' Last Exam (CLI) | 23.8 | 28.5 | Beats Kimi K3 & Opus 4.8 |
| CyberGym (vuln discovery) | 77.2 | 84.5 | Global #1 (>GPT-5.6 Sol 83.6) |
| ExploitBench | 24.4 | 54.4 | 2x+ |
| SWE-Marathon v1.1 | 19.4 | 42.5 | ~2x |
| AutomationBench | 26.2 | 48.2 | — |
| FrontierSWE | 67.5 | 78.1 | — |
| GDPval-AA v2 (Elo) | 1508 | 1769 | Highest in table |
Artificial Analysis Intelligence Index: 60 — the frontier band alongside Claude Fable 5 and GPT-5.6 Sol, and tied with Kimi K3 for the top open-source score.
Why "Post-Training Only" Matters
GLM-5.3 is the cleanest proof that training technique beats raw scale. No new architecture, no added parameters — just a massively scaled post-training pipeline built on an "AI researcher writes questions, AI grader verifies" synthetic loop, with long-horizon coding training sped up 2.3x. Chinese labs are clearly shifting from "scale at any cost" to "squeeze more from less."
And then there's the security surprise: before launch, a red-team review with Tsinghua and Nankai found 2,436 vulnerabilities in 269 open-source projects (1,097 medium-to-high severity) — including a 40-year-old DNS protocol-level flaw with 10M+ public DNS services potentially exposed. That capability is exactly why the weights are delayed to Aug 28 for a security review.
Pros & Cons
Pros
- ✅ AA 60 — tied for the top open-source Intelligence Index
- ✅ Open-source #1 coding; ~50% better than 5.2
- ✅ CyberGym global #1 — unexpected security capability
- ✅ Outstanding token efficiency, lowest frontier single-task cost
- ✅ 1M context + MIT open weights (Aug 28)
- ✅ Price unchanged in a market of hikes
Cons
- ❌ Weights not out until Aug 28 — can't self-host yet
- ❌ Text-only, no multimodal
- ❌ Stock dropped on launch day (lukewarm market reaction)
- ❌ Some figures are Zhipu-internal until weights are verified
- ❌ Hardest reasoning tasks still trail closed frontier models
Who Should Use It (And Who Shouldn't)
Great fit
- Developers: long-horizon coding, terminal automation, codebase-level tasks
- Agent builders: multi-step tool calls, real software engineering
- Security researchers: vulnerability discovery and red-team work
- Cost-conscious teams wanting near-frontier open models
Poor fit
- Multimodal applications (text-only)
- Strict private deployment before Aug 28 weights
- Creative / long-form writing (not its strength)
Frequently Asked Questions
Is GLM-5.3 open source?
Weights release ~Aug 28 under MIT on Hugging Face (zai-org), delayed for security review. API is live now.
How big is it?
743B total MoE params, ~40B active, same base as GLM-5.2 — all gains from post-training.
How much does it cost?
$1.40/M input, $4.40/M output — unchanged from 5.2; lowest single-task cost among frontier flagships.
Is it really the best open coding model?
Terminal-Bench 28.3 is open-source #1; DeepSWE 66.9 beats DeepSeek V4 Pro; internal coding ~50% better than 5.2.
Is the cybersecurity ability real?
CyberGym 84.5% global #1; red-team found 2,436 vulns in 269 projects, incl. a 40-year DNS flaw affecting 10M+ services.
What is the context window?
1M input / 128K output, text-only.
How does it compare to Claude Fable 5 / GPT-5.6?
AA Index 60 puts it in the frontier band; it ties Kimi K3 for top open score and is far more token-efficient.
Sources
- 觉醒AI — GLM-5.3 API live: AA 60, tied open-source #1, lowest frontier task cost
- 36氪 — 智谱 VS DeepSeek: two forked paths of self-evolution
- Aliyun dev — GLM-5.3: post-training only, coding +50%, cyber security emerges
- 量子位 — PhanRouter launches on GLM-5.3
- apidog — What is GLM-5.3? Zhipu's open-weight coding model
- orcarouter — GLM-5.3 vs GLM-5.2: Same Brain, Doubled Benchmarks
没有评论:
发表评论