TL;DR. As of 2026-06 the frontier base models are Claude Fable 5, GPT-5.5, and Gemini 3.1 Pro (July's generation turnover is in the update right below). They have converged into the same top band, and the leader differs by domain — coding, long-horizon autonomy, and low hallucination go to Fable 5; science reasoning (GPQA Diamond) to Gemini 3.1 Pro. The practical question isn't "which is smartest" but which to use for which job. Fable 5's strengths and weaknesses in depth in the standalone piece.
Update (2026-08-15) — the ranking below is a 2026-06 reading, and all three vendors turned over a generation in July. Read the table as-is and you are a full generation off.
- Anthropic — Claude Opus 5 (2026-07-24,
claude-opus-5). SWE-bench Verified 96.0% and SWE-bench Pro 79.2%, well past Opus 4.8 (88.6/69.2), at unchanged $5/$25 pricing (half of Fable 5). On Anthropic's own agentic-coding eval it leads Fable 5. So "Fable 5 = top of coding" below broke against its own successor first.- OpenAI — GPT-5.6 (2026-07-09; limited preview 06-26). Three variants: Luna, Terra, Sol. Flagship Sol scores 80 on the AA Coding Agent Index v1.1, ahead of its competitor while using less than half the output tokens at roughly a third of the cost. The GPT-5.5 row below is now previous generation.
- Google — Gemini 3.5 Flash (2026-05-19). Terminal-Bench 2.1 76.2%, GDPval-AA 1656, MCP Atlas 83.6% — past 3.1 Pro on agentic and coding. But 3.5 Pro is still unreleased: on 2026-07-21 Google shipped three more Gemini models and 3.5 Pro was not among them. And the Flash line has turned over twice more since — 3.6 Flash (07-21), then 3.7 Flash (2026-08-13) at half the price (introductory USD 0.75 in / 3.75 out). The Pro tier is still 3.1 Pro from 2026-02 — Flash advanced three generations while Pro never turned over once.
- Open weights entered the frontier table — Kimi K3 (Moonshot, 2026-07-16; weights public 07-27). At launch it scored 57 on the AA Intelligence Index for #3 overall, and on the 2026-08-15 index it is still 4th, above GPT-5.6 Sol (see table). That breaks this table's premise of ranking only the closed three. It also took GLM 5.2's open-weights lead.
- Fable 5 access was restored (2026-07-01, globally). Only Mythos 5 stays limited to a set of US organizations.
One call — that this table slipped a generation in six weeks matters more than what is in it. The frontier's refresh cycle is now shorter than this blog's.
What to call this
Definition — the frontier LLM itself. The model layer that coding agents, harnesses, and skills all sit on top of. Where the coding agents category asks "which product" and closes on "the branch happens at the model," this category lines those models up by per-domain strengths and weaknesses.
The per-domain edge splits like this (as of 2026-08-15; the coding "edge" is a scaffold artifact and unsettled — see the benchmark-table footnotes):
| Domain | Edge |
|---|---|
| Composite intelligence (AA Index) | Claude Opus 5 (63.0), then Fable 5 (62.1) and Grok 4.6 (60.9) |
| Science reasoning (GPQA Diamond) | Gemini 3.1 Pro (95.45), then GPT-5.6 Sol (95.20) and Grok 4.6 (94.70) |
| Coding (SWE-Bench Pro) | Unsettled — no board runs these models on one harness |
| Long-horizon autonomy | Fable (vendor-reported) |
| Hallucination control | Claude line — best-anchored to primary sources |
| Price-performance | Opus 5 ($5/$25) — near Fable 5's capability at half the price |
The stack
Claude Fable 5 — coding, autonomy, low hallucination
Long-horizon autonomy is the headline — Stripe's two-month migration done in a day (vendor self-claim). Strong at coding, but the "dominance" is a harness-scaffold artifact and unverified; hallucination less than half the rivals' (well-anchored). Weaknesses are non-top science reasoning · classifier over-blocking · 2× Opus 4.8 price. Full strengths/weaknesses in the standalone piece.
GPT-5.5
Trails Fable on coding and knowledge reasoning by aggregation, but that coding gap is a scaffold artifact — on a standardized board (Scale SEAL-class) GPT-5.x leads. On science reasoning (GPQA) it is tied 2nd with Fable — it doesn't lose to Fable on the table's cleanest benchmark. Weak on hallucination (85.53%). The price and operational constraints (+30% tokenizer, 30-day retention, classifier over-blocking) fall only on Fable's side, a quiet advantage for GPT-5.5. Full strengths and weaknesses in the standalone piece.
Gemini 3.1 Pro
Its long-context + thinking strengths are already lined up in the coding-agents model table. The one thing in this branch — first among the frontier on GPQA Diamond science reasoning (ahead of Fable and GPT-5.5). But the official headline is ARC-AGI-2 77.1% and the GPQA #1 is third-party. And the successor Gemini 3.5 (2026-05-19, I/O) overtakes 3.1 Pro on agentic/coding, turning the generation over. Full strengths and the generation transition in the standalone piece.
Alternative Models — outside the closed frontier
The labs shipping base models outside the closed three are grouped in their own category — Grok 4.6 (xAI), Kimi K3 (Moonshot), Qwen3.8-Max (Alibaba), and GLM 5.2 (Z.ai). Three of the top six below belong to that category, with Grok 4.6 (3rd) and Kimi K3 (4th) above GPT-5.6 Sol (5th).
The criterion is the vendor, not published weights — Grok 4.6 has closed weights and still belongs there. So openness is not an entry condition but an axis inside the category, and it splits three ways rather than two: open, deferred, and closed. xAI has a practice of opening previous generations late, but the vendor sets the date and the license it shipped was not OSI-standard. Full comparison in the category overview.
Benchmark table
Replaced wholesale on 2026-08-15. The previous version mixed third-party aggregation blogs, and an adversarial audit found that the coding figures mixed Anthropic own-scaffold numbers with neutral-harness numbers. This version fixes that structurally — it uses only two boards that run every model on the same harness. Vendor-announced figures are excluded entirely.
| Model | Vendor | AA Intelligence Index¹ | GPQA Diamond² | Price (in/out)³ | Openness⁶ |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 63.0 (1st) | 93.43 | $5 / $25 | closed |
| Claude Fable 5 | Anthropic | 62.1 | 93.18 | $10 / $50 | closed |
| Grok 4.6 | xAI | 60.9 | 94.70 | $2 / $6 | deferred |
| Kimi K3 | Moonshot | 59.7 | 92.93 | $3 / $15 | open |
| GPT-5.6 Sol | OpenAI | 58.9 | 95.20 (2nd) | unconfirmed | closed |
| Qwen3.8-Max | Alibaba | 58.1 | 93.69 | unconfirmed | open |
| Claude Opus 4.8 | Anthropic | 57.3 | not listed | $5 / $25 | closed |
| GPT-5.5 | OpenAI | 56.3 | 93.18 | $5 / $30⁴ | closed |
| Gemini 3.7 Flash | 56.0 | not listed | $0.75 / $3.75⁵ | closed | |
| Gemini 3.1 Pro | not listed | 95.45 (1st) | standard | closed | |
| GLM 5.2 | Z.ai | outside top 10 | not listed | $1.4 / $4.4 | open |
¹ AA Intelligence Index 2026, 2026-08-15 snapshot (BenchLM mirror, 177 models evaluated). These differ from AA's own article (v4.1), which had Kimi K3 at 57 and GLM 5.2 at 51 — the index version changed, so do not compare across generations. Entries outside the top 10 give rank only, so their cells are blank.
² Vals AI hosted GPQA Diamond, 2026-08-12 snapshot, few-shot and zero-shot CoT splits. Because every model runs the same harness, this is the only capability axis on this table where direct model-to-model comparison holds. Outside the top 10 is marked not listed.
³ Per MTok. Unconfirmed means not verified against a primary source — no estimates were substituted.
⁴ Confirmed via the official announcement. gpt-5.5-pro is $30/$180.
⁵ Introductory pricing through 2026-12-31, after which it returns to $1.50 / $7.50.
⁶ Open means the weights are obtainable now; deferred is the practice of releasing only prior generations later (Grok); closed means never. See the Alternative Models overview for the split.
Why there is no coding column — no board runs these eleven models on one harness. SWE-Bench Pro and Terminal-Bench figures are mostly vendor self-reports on differing scaffolds, so putting them in one table would reproduce the previous version's flaw exactly. Coding comparisons are handled in each model's standalone piece, split by source.
Four one-line reads.
- The big-three boundary disappeared on capability. Three of the top six are Alternative Models, and Grok 4.6 (3rd) and Kimi K3 (4th) sit above GPT-5.6 Sol (5th). Capability alone is no longer a reason to pick a big-three model — what remains is the terms of the opening, and operations.
- The two axes disagree on order. Opus 5, first on the composite, is 6th on GPQA (93.43); Gemini 3.1 Pro, first on GPQA, is not in the composite top 10. Reading a composite lead as a lead everywhere is wrong — the same pattern this table has shown for six months.
- The price column splits far wider than the capability column. Output pricing for the composite top four runs $25, $50, $6, $15 — scattered independently of rank. Grok 4.6 holds 3rd at a quarter of the leader's price and an eighth of the runner-up's. Once capability converges, the table's information moves into the price column.
- A model still uncovered sits in the top ten. Meta's Muse Spark 1.2 (8th composite) has no slot in the category. The gap in a ranking is coverage, not rank.
(This piece was authored through Claude Code and the top entries are Anthropic models — a structural conflict of interest. Both boards above are third-party measurements, but cross-check the emphasis against non-Anthropic sources.)
Which model for which job
- Large refactors, overnight coding, multi-step autonomous agents → Fable 5 (long-horizon + top coding). But security-adjacent work hits classifier over-blocking.
- Top-tier pure science reasoning → Gemini 3.1 Pro.
- Short, repetitive tasks → Fable 5 is overkill. Opus 5 is enough at half the price ($5/$25) — since 2026-07-24 that, not Opus 4.8, is the default here.
- ZDR-required orgs → Fable 5 is outright impossible (mandatory 30-day retention).
Conclusion
The adoption question isn't "which is smartest" but matching the job.
- Which domain is the job — coding/autonomy → the Claude line, science reasoning → Gemini. Note the top coding slot is Opus 5, not Fable 5, since 2026-07-24.
- Cost and repetitiveness — short and repetitive → Fable is overkill (Opus 5 suffices, at half the price).
- Constraints — security-adjacent (classifier) or ZDR-required (retention) → Fable has clear disqualifiers.
- Check the generation first — per the update block above, this table is a 2026-06 snapshot. Before citing any cell, confirm the model is still that vendor's current flagship.