leejk/ jk lee
Models: lining up the frontier base models

Landscape · 2026.08

Alternative Models — base models outside the closed frontier

Once four labs outside the big three climbed into the frontier band, the real fork moved down to the terms of the opening and to price

·
#models#open-weights#glm#kimi#qwen#grok#landscape

The labs shipping base models outside Claude, OpenAI, and Google, gathered in one slot. Kimi K3 (Moonshot), Qwen3.8-Max (Alibaba), GLM 5.2 (Z.ai), and Grok 4.6 (xAI) arrived back to back between June and August 2026, and on the AA Intelligence Index they now sit in the same top band as the closed frontier. One call — what decides adoption here is not the benchmark score but the terms of the opening, and openness is not a binary but a three-way split: open, deferred, and closed. GLM is cleanest at MIT, Kimi K3 ships a permissive house license, Qwen3.8-Max publishes a checkpoint narrower than its API model, and Grok 4.6 is closed today but comes from a lab that opens previous generations late.

TL;DR. The slot for labs shipping base models outside the closed frontier (Claude, OpenAI, Google). Four arrived back to back between June and August 2026 — Kimi K3 (Moonshot, 07-16), Qwen3.8-Max (Alibaba, 08-03), GLM 5.2 (Z.ai, 06-14), and Grok 4.6 (xAI, 08-12). Capability is already the same bandthree of this category sit in the AA Intelligence Index top six, two of them above GPT-5.6 Sol. One call — the fork is not the benchmark but the terms of the opening, and those split three ways: open, deferred, and closed.

What this slot is

Definition — the vendor column in models for base models belonging to none of the three closed frontier labs (Anthropic, OpenAI, Google). The criterion is the vendor, not whether weights are publishedGrok 4.6 has closed weights and still belongs here. Openness is not an entry condition but the axis that separates models inside this category.

Why one slot rather than one page per lab — there are four labs, and four vendor pages would leave one model in each. The axes that actually separate them are the terms of the opening and price, so they belong on one table.

The stack

Model Lab Openness License AA Intelligence Index Price (in/out)
Kimi K3 Moonshot AI Open (07-27) Kimi K3 License (house, permissive) 59.7 (#4 overall) $3 / $15
Qwen3.8-Max Alibaba Open (08-12) qwen3.8-max (house, not OSI) 58.1 (#6 overall) unconfirmed
GLM 5.2 Z.ai (formerly Zhipu) Open (06-14) MIT outside top 10 $1.4 / $4.4
Grok 4.6 xAI Deferred proprietary (prior generations opened later) 60.9 (#3 overall) $2 / $6

Index figures are the 2026-08-15 snapshot. They are a different index version from AA's own article (Kimi K3 57, GLM 5.2 51), so they are not used for cross-generation comparison. Three of the four sit inside the overall top six, and Grok 4.6 is #3 while being the cheapest on this table.

Each model's strengths and weaknesses are sourced in its standalone piece.

  • Grok 4.6 — holds #3 overall while being the cheapest in the frontier band ($6 output). But the weights are closed and context is 500K, half the field.
  • Kimi K3 — 2.8T MoE, 1M context, native vision. The highest capability among the open ones and the largest open weights published. Also the most expensive of the three open models.
  • Qwen3.8-Max — a 2.4T MoE that opened the Max tier for the first time. But its published checkpoint differs from the API model.
  • GLM 5.2 — lower on capability but MIT-licensed and the cheapest. Still the slot when you want licensing risk at zero.

The real fork is the terms of the opening, not capability

All four demo fine. Three axes decide adoption.

Axis 1 — openness is not a binary but three positions. Asking "open or closed" fails to separate this category. There are three.

Position Models Can you get the weights now What it means for adoption
Open Kimi K3, Qwen3.8-Max, GLM 5.2 Yes Revocation is structurally impossible
Deferred Grok 4.6 No; the previous generation may open when the next ships Cannot be used as the basis of a plan
Closed Claude, OpenAI, Google No Access is bound to the vendor and regulators

The middle term matters. xAI did open Grok-1 under Apache 2.0 in 2024-03 and the Grok 2 line in 2025 — the practice is real. But the vendor sets the date, and the license attached when it opened was not OSI-standard (Grok 2.5 in 2025-08 came under a house license with anti-competitive terms). It will open eventually is not a basis for an adoption plan.

Axis 2 — even among the open three, licensing splits three ways. The premise that "open weights" is one thing breaks here. GLM 5.2 is MIT (HF card) — standard OSI, so legal review is effectively done. Kimi K3 ships the Kimi K3 License, a permissive modified-MIT-style house license that is not OSI-approved. Qwen3.8-Max carries a house license named qwen3.8-max (HF card). All three let you take the weights; what you may then do with them is not the same. In enterprise adoption this binds before a three-point benchmark gap does.

Axis 3 — the published checkpoint is not the API model. Qwen3.8-Max surfaced this. The API's Qwen3.8-Max was introduced with text, image, and video input at 1M context, yet the published Qwen/Qwen3.8-2.4T-A95B card states it is text-only ("Multimodal inputs are not supported"), with context 262,144 native, extensible to ~1,010,000. The model you picked off a benchmark may not be the model you downloaded. Reading "weights published" as "you get the same thing" is wrong — cross-checking the card's modality and context against the API spec is now a required step here.

Price is the second substantive axis

Once capability converges, price is what remains — and this category took that axis. Line up the output prices of the composite top four in the models table and it splits: Fable 5 $50, Opus 5 $25, Kimi K3 $15, Grok 4.6 $6. The #3 model's output price is a quarter of the leader's, and GLM 5.2 is cheaper still at $4.4.

But measure cost per task, not per token. GLM 5.2 is verbose, landing at $0.32 per task on AA against Kimi K3's $0.94 — a roughly 3x gap in cost per task against a roughly 2x gap in token pricing.

Why open this category now

Two signals landed within eight weeks.

First, access to the closed frontier was proven not to belong to the vendor. Claude Fable 5 had its access cut three days after launch by a US export-control directive (2026-06-12). It was lifted on 07-01, but the switch was never the vendor's to hold.

Second, the capability gap stopped being an adoption blocker. Three of this category sit in the AA top six, with Grok 4.6 (#3) and Kimi K3 (#4) above GPT-5.6 Sol (#5). That breaks the models table's premise of ranking only the closed three.

Conclusion

The adoption question is no longer "is anything outside the big three good enough." That closed in the summer of 2026. Four questions remain.

  1. Can you absorb losing access — if not, pick from the open three. Treat deferred opening (Grok) exactly as you treat closed.
  2. How much license can you carry — MIT ends the review (GLM). A house license means reading the clauses yourself (Kimi, Qwen).
  3. Is the checkpoint you get the model from the benchmark — verify modality and context on the card. Qwen3.8-Max differs.
  4. Did you price it per task — not per token, per task. GLM $0.32, Kimi K3 $0.94.

Sources

Sub-documents