GLM-5.2 Is The New Best Open Model
open-modelsbenchmarksglmchinese-aimodel-evaluation
Abstraction: GLM-5.2 open model capabilities review, benchmarks
Key points:
- GLM-5.2 (from Z.ai/Zhipu) is the strongest open model, roughly at Opus 4.7 level on text-only tasks; Zvi estimates it 4-7 months behind the frontier, closing the gap more than DeepSeek R1 did at its peak.
- Benchmarks: Artificial Analysis v4.1 = 51 (behind Fable 60, Opus 4.8 56, GPT-5.5 55, Opus 4.7 54); LiveBench between Opus 4.5-4.6; #1 on PosttrainBench slightly ahead of Opus 4.8; API cost $1.40/$0.26/$4.40 per M tokens, token-hungry so effectively pricey for an open model.
- Strongly evidenced to be distilled from Claude Opus (identifies as Claude, "Claude voice," Claude harness); distillation means it overperforms on benchmarks/common tasks and generalizes poorly, causing underestimation of the true frontier gap.
- Mixed reception: Jeremy Howard says "at least as good as Opus 4.8 and GPT 5.5"; critics (QC, gwern) call it "benchmaxxed," weak outside coding, no native vision, sycophantic. Best at "puzzly" coding, weaker on real conversational/messy tasks.
- Z.ai founder Jie Tang claims a Mythos-level model this year; Musk speculates Q1 2027; Zvi would bet against Fable-5 parity by EOY 2026 but not by Q2 2027.
Connections: Zvi Mowshowitz · Zhipu AI · Claude · Open Weight Models · Model Distillation · Model Evaluation
Source: https://thezvi.substack.com/p/glm-52-is-the-new-best-open-model