TL;DR. Grok 4.6 (xAI, 2026-08-12) holds one slot: the cheapest model in the frontier band. Strengths — 60.9 on the AA Intelligence Index for #3 overall at $2 input / $6 output, a quarter of the output price of Opus 5 ($5/$25) directly above it. Also 94.70 on Vals GPQA Diamond for #3 in science reasoning. Weaknesses — the weights are closed, context is 500K, half the field, and the parameter count is undisclosed. xAI does have a practice of opening previous generations late, but the license it actually shipped was not OSI-standard.
Provenance note. This is not a hands-on review of Grok 4.6. It is read off the BenchLM model page (specs, pricing, weights status) and the AA Intelligence Index and Vals GPQA Diamond snapshots. The generational deltas (DeepSWE, APEX-Agents) come from third-party coverage relaying the announcement and are marked as author-reported. No measured or felt observations.
At a glance
| Item | Value |
|---|---|
| Model | Grok 4.6 (xAI) |
| Release | 2026-08-12 |
| Context | 500K tokens (knowledge cutoff 2026-02-01) |
| Modality | Text + image input → text output |
| Reasoning control | Four levels: low, medium, high, xhigh |
| Price | $2 / $0.50 / $6 per MTok (input / cache hit / output). The fast variant is 2x |
| Parameters | Undisclosed (provider does not publish, per BenchLM) |
| Weights | Closed — not published, proprietary |
| Access | xAI API, Grok Build, Cursor, OpenRouter, Vercel, Cloudflare |
Why it belongs here — the third position on openness
This category is the slot for base models outside the closed frontier three (Claude, OpenAI, Google). Grok 4.6 meets that condition, but its weights are closed — it sits in the same column as Kimi K3, Qwen3.8-Max, and GLM 5.2 while standing on the opposite side of the openness axis.
Yet xAI is not permanently closed either. It has a practice of opening the previous frontier when the next generation ships — Grok-1 under Apache 2.0 in 2024-03, and the Grok 2 line's weights in 2025. So the openness axis in this category is not a binary but a three-way split.
| Position | Models | Can you get the weights now |
|---|---|---|
| Open | Kimi K3, Qwen3.8-Max, GLM 5.2 | Yes |
| Deferred | Grok 4.6 (xAI) | Not now; the previous generation may open when the next ships |
| Closed | Claude, OpenAI, Google | No |
That is what including Grok buys this category — a model outside the big three yet closed exposes the middle term of the openness axis.
Strengths — what it does well
1. Cheapest in the frontier band — the substance of this model
It sits at 60.9 on the AA Intelligence Index for #3 overall (2026-08-15 snapshot) while priced at $2 input / $6 output. Against the rest of that table's top, the gap is wide.
| Model | AA Index | Output price |
|---|---|---|
| Claude Opus 5 | 63.0 | $25 |
| Claude Fable 5 | 62.1 | $50 |
| Grok 4.6 | 60.9 | $6 |
| Kimi K3 | 59.7 | $15 |
It holds #3 overall at a quarter of the leader's output price and an eighth of the runner-up's. Cache-hit input drops to $0.50. That makes a real trade available — give up 2.1 index points, cut output cost fourfold — and in high-volume calls and long-running agents that axis becomes effective cost immediately.
2. Third in science reasoning — on the same harness
It scores 94.70 on Vals AI hosted GPQA Diamond (2026-08-12 snapshot), behind Gemini 3.1 Pro (95.45) and GPT-5.6 Sol (95.20), and ahead of Opus 5 (93.43) and Fable 5 (93.18). Because every model runs the same harness, this is an axis where direct comparison holds in the models table — and the contrast sharpens once price is read alongside it.
3. Large generational gains in agentic and coding work (author-reported)
Against Grok 4.5, the announcement reports DeepSWE 54 → 65.9 and APEX-Agents 47.1 → 57.5. xAI frames this generation around long-running agents, agentic coding, and interactive work. Aggregate boards echo the emphasis — per BenchLM, agentic (66.7) ranks ahead of coding (63.6). The generational deltas are vendor figures, so read direction until a standardized board.
4. Four levels of reasoning effort
Reasoning effort is selectable across low, medium, high, and xhigh. For giving each task its own token budget, that is a wider dial than models offering a binary.
Weaknesses — what it does poorly
1. The weights are closed
This is the decisive split from the other three in this category. Per BenchLM, weights not published, proprietary. The access-revocation risk Fable 5 demonstrated on 2026-06-12 applies here in full — being cheap only holds while the vendor keeps it open.
2. Deferred opening is a practice, not a commitment — and what shipped was not OSI
Two caveats attach to xAI's opening practice. First, there is no guarantee of timing — release is tied to the next generation shipping, and the vendor sets that date. Second, the license attached when it did open was not OSI-standard: Grok-1 was Apache 2.0, but Grok 2.5, released in 2025-08, came under a house license with anti-competitive terms. Placing that on the same axis as GLM 5.2's MIT is wrong. Do not build an adoption plan on it will open eventually.
3. 500K context — half the field
Most of this category and the frontier top run 1M; Grok 4.6 runs 500K. Kimi K3, GLM 5.2, Opus 5, and Gemini 3.7 Flash are all at 1M or above. That constraint collides head-on with the long-running agents positioning, and it binds first on workloads that hold a whole large repository in one pass.
4. The parameter count is undisclosed
Per BenchLM, the provider does not publish parameters. Some coverage prints 1.5T, but there is no vendor confirmation. Practical impact is small since self-hosting was never on the table, but it contrasts with the other three in this category, which publish both total and active counts — degree of openness is not only about weights.
5. Knowledge cutoff of 2026-02
The cutoff is 2026-02-01, half a year before the 08-12 release. The 2026 first-half shifts this blog covers — the Fable 5 episode, GPT-5.6, three open-weight releases — are not inside the model. For anything that leans on the model's own ecosystem knowledge, search or context injection is mandatory.
What to use it for
- Frontier-grade work under cost pressure → Grok 4.6. Holding #3 overall at the lowest price is a slot it occupies alone in this table.
- High-volume calls and long-running agents → Grok 4.6, but confirm first that the work fits in 500K.
- You need 1M context → Kimi K3, GLM 5.2, or Opus 5.
- You must structurally eliminate revocation risk → not this model. Pick from the open three.
- Top science reasoning → Gemini 3.1 Pro (95.45). Grok is close at #3 (94.70) but not first.
Conclusion
Grok 4.6 does not win on capability; it reset the price axis.
- Price — #3 overall at a quarter and an eighth of the output price above it. A slot it holds alone in this table, and the firmest claim in this piece (both boards are third-party measurements).
- Openness — closed now; deferred opening is a practice, not a commitment; and what shipped was not an OSI license. It does not belong on the same axis as the other three here.
- Constraints — the 500K context and the 2026-02 cutoff collide with its own long-running agents positioning. Measure both before picking it for being cheap.
Sources
Aggregate and independent evaluation
- BenchLM — Grok 4.6 model page (specs, pricing, weights status, aggregate ranks)
- AA Intelligence Index 2026 snapshot (#3 overall, 60.9)
- Vals AI hosted GPQA Diamond snapshot (#3 science reasoning, 94.70)
Opening practice
- xAI — Open Release of Grok-1 (2024-03, Apache 2.0)
- TechCrunch — xAI open-sources Grok 2.5 (2025-08, house license with anti-competitive terms)