leejk/ jk lee

Alternative Models — base models outside the closed frontier · 2026.06

GLM 5.2

Strong on its 2026-06 open-weights intelligence lead, MIT license, 1M context, and 1/10th the price; weak on token verbosity, self-claimed benchmarks, and the limits of a single composite lead — it lost the lead to Kimi K3 in July

·
#models#zai#zhipu#glm#open-weights#landscape

What GLM 5.2 (Z.ai/Zhipu) does well and does poorly, as strengths and weaknesses. Strengths — leading open-weights model on the Artificial Analysis Intelligence Index v4.1 at 51 (vs MiniMax-M3 / DeepSeek V4 Pro max at 44, Kimi K2.6 at 43), edges GPT-5.5 xhigh (1514) on one aggregate (GDPval-AA v2 at 1524), MIT license with 1M context, input $1.4 / output $4.4 at roughly 1/10th of Fable 5, and open weights mean access cannot be revoked. Weaknesses — 43k output tokens per task, more verbose than rivals (24-37k), so the price edge is partly clawed back; the coding and reasoning benchmarks (SWE-bench Pro 62.1, AIME 99.2, GPQA 91.2) are vendor self-claims with only the AA index independently verified; and a single composite lead is not a lead everywhere (GPQA trails Gemini, Fable, and GPT). Not hands-on — read off the independent index, the HF card, and the announcement, split by source.

TL;DR. GLM 5.2 (zai-org/GLM-5.2, Z.ai / formerly Zhipu) is the leading open-weights base model. Strengths — top of the Artificial Analysis Intelligence Index v4.1 at 51 among open models (vs MiniMax-M3 / DeepSeek V4 Pro max at 44, Kimi K2.6 at 43), edges GPT-5.5 xhigh (1514) on one aggregate (GDPval-AA v2 at 1524), and ships MIT-licensed with a 1M context at ~1/10th of Fable 5's priceopen weights mean access can't be revoked. Weaknesses — at 43k output tokens per task it is more verbose than rival open models (24-37k), clawing back part of the price edge; its coding and reasoning benchmarks are vendor self-claims with only the AA index independently verified; and a composite lead is not a lead everywhere (GPQA trails Gemini, Fable, and GPT). Read the capability signals by source.

Update (2026-07-27) — the open-weights lead passed to Kimi K3. The "open-weights #1" below is a 2026-06 reading. Moonshot AI's Kimi K3 (2.8T MoE, 1M context, launched 2026-07-16) scored 57 on the AA Intelligence Index for #3 overall — six points ahead of GLM 5.2 (51) — and its weights went public on July 27 under the Kimi K3 License, taking the open-weights lead. But the spine of the analysis below holds: GLM 5.2 is still far cheaper (AA cost per task $0.32 vs K3's $0.94) and MIT-licensed, which is cleaner than the bespoke Kimi K3 License (open weights, but not OSI open source). So: the intelligence lead is gone; the cost and licensing axes remain.

Provenance note. This is not a hands-on review of GLM 5.2. It is read off AA's independent intelligence index, the HF model card, and the announcement, with numbers split between independent measurement (AA ran it) and author-reported (HF card self-claims). No measured or felt observations.

At a glance

Item Value
Model ID zai-org/GLM-5.2 (Z.ai, formerly Zhipu AI)
Architecture MoE — 744B total / 40B active (AA aggregate; HF card says ~753B total), sparse attention
Context 1M tokens (up from 200K on GLM-5.1)
License MIT (open weights, commercial use)
Price $1.4 / $0.26 / $4.4 per MTok (input / cache hit / output) — ~1/10th of Fable 5 ($10/$50)
Release Shipped 2026-06-14 → AA open-weights #1 on 06-18
Availability Z.ai first-party API + third parties (DeepInfra, Fireworks) and self-host

Why now — an open-weights #1 where the closed frontier just wobbled

The models overview closed by saying the #1 base model splits by domain — among the three closed frontier labs (Fable, GPT, Gemini). GLM 5.2 enters from outside that frame. Right after launch it topped the AA Intelligence Index among open models, and at the same moment the signals converged — 826 points on HN, past 1,000 likes on HF.

Timing is the point. Claude Fable 5 had its access cut to all foreign nationals three days post-launch by a US export-control directive (2026-06-12; lifted 07-01). That week proved a closed base model can be revoked on geopolitics regardless of capability — it came back 18 days later, but the switch was never the vendor's to hold. GLM 5.2 walked into exactly that spot with MIT weights — climbing into the top capability band as a model whose access itself cannot be revoked. On top of "which one is smartest," the live question is now "how far has the one nobody can switch off gotten?"

Strengths — what it does well

1. Leading open-weights intelligence — the one claim with independent backing

On the AA Intelligence Index v4.1 it scores 51, ahead of MiniMax-M3 / DeepSeek V4 Pro max (44 each) and Kimi K2.6 (43). This is a composite index AA ran itself, making it the best-sourced capability line in this piece.

2. Edges GPT-5.5 on one aggregate

On GDPval-AA v2 it scores 1524, narrowly above GPT-5.5 xhigh (1514) (AA). An open-weights model passing a closed frontier model on one aggregate is the headline number. But GDPval-AA is a single Elo aggregate and the gap is 10 points — don't generalize it to "smarter than GPT-5.5." It backs only that it led in one place.

3. Price and license — non-revocable and 1/10th the cost

It is MIT-licensed (open weights, commercial use). Price is $1.4 input / $4.4 output — about 1/10th of Fable 5 ($10/$50), with cost-per-task ~$0.46 (AA). Take the weights and self-host, and no vendor can cut your access — the revocation risk Fable 5 proved on 2026-06-12 (lifted 07-01) structurally doesn't apply. For orgs with hard ZDR or data-governance constraints, this axis alone makes it a candidate.

4. 1M context + long-horizon work (author-reported)

Context jumped from GLM-5.1's 200K to 1M. The HF card claims it "stably sustains 1M tokens for long-horizon work" and lists coding with flexible effort, tool use, and agentic tasks (HF card). This long-horizon claim is author-reported.

5. Author-reported coding and reasoning benchmarks — strong but self-claimed

Per the HF card — AIME 2026 99.2, GPQA-Diamond 91.2, HLE 40.5; SWE-bench Pro 62.1, Terminal Bench 2.1 82.7; MCP-Atlas 76.8 (HF card). Community posts also circulated that GLM beat GPT on SWELancer. The direction (strong coding/agentic) is consistent, but these numbers are author-reported, not standard-board entries — read split, below.

Weaknesses — what it does poorly

1. Token verbosity — claws back part of the price edge

The most practical downside. AA measures 43k output tokens per index task (37k reasoning)more than MiniMax-M3 (24k) and DeepSeek V4 Pro max (37k) (AA). Cheap per token plus many tokens per task raises effective cost and latency. 1/10th the price is not 1/10th the cost — verbosity eats into it.

2. Only one line is independently verified

The one solidly verified figure is the AA composite index (51, open-weights #1). The detailed coding/reasoning benchmarks (SWE-bench Pro 62.1, etc.) are all author-reported, with no standard-board (Scale SEAL class) comparison where every model runs the same harness. Same trap seen on Fable 5self-claimed coding numbers can be scaffold artifacts. Hold detailed-superiority claims until an independent board backs them.

3. A composite lead is not a lead everywhere

GPQA-Diamond 91.2 (author-reported) trails the closed frontier — Gemini 3.1 Pro ~95 (science-reasoning #1), Fable 5 / GPT-5.5 ~93. So it is open-weights #1 overall, not the top on pure science reasoning. Reading "index #1" as superiority everywhere is wrong.

4. Governance — weights are open, but the source is a Chinese lab

Data, content-policy, and censorship alignment are a separate axis to evaluate when adopting a Chinese-lab model. The mitigant: MIT-open weights let you self-host and cut external API dependence and telemetry — a trade-off versus the closed frontier, not a disqualifier.

Which job to use it for

  • Avoiding access-revocation risk, self-host, hard ZDR → GLM 5.2. Holding the weights means no vendor or geopolitics can cut access (Fable 5 got cut).
  • Cost-sensitive bulk inference → GLM 5.2 (~1/10th per token). But fold verbosity (43k/task) into the cost model.
  • Top-tier pure science reasoning → still Gemini 3.1 Pro (GPQA). GLM 5.2 leads the composite, not this domain.
  • Verified top-tier coding → still self-claimed. Hold the verdict until a standard board, and re-rank on the same axis as Fable 5 and GPT-5.5.

Conclusion

The real news in GLM 5.2 is not one benchmark cell — it is that an axis was added.

  1. Capability — the open-weights composite lead was independently backed. But that is a 2026-06 reading, ceded to Kimi K3 (57) on 07-27, and detailed coding/reasoning superiority is self-claimed; hold until a standard board.
  2. Cost — 1/10th per token but verbose, so recompute effective cost on tokens-per-task. This axis holds against Kimi K3 too (AA cost per task $0.32 vs $0.94).
  3. Non-revocability — MIT weights mean access can't be cut. The risk the closed frontier proved on 2026-06-12 structurally doesn't apply — and that is as large an adoption variable as the capability ranking. That risk was reversed on 07-01, but the reversal came from the regulator rather than the vendor, which hardens the axis rather than softening it.

Sources

Primary / independent evaluation

Same topic