TL;DR. Claude Fable 5 (
claude-fable-5) is the most capable model Anthropic has released generally. Strengths — long-horizon autonomy, coding, vision, memory use, and low hallucination + reasoning-token efficiency. Weaknesses — it is not the sole leader on science reasoning (behind Gemini and GPT-5.5 on GPQA Diamond), its safety classifiers over-block security and life-sciences-adjacent work, and at 2× Opus 4.8's price it is overkill for repetitive tasks. Operational constraints follow too (mandatory 30-day retention, no ZDR, +30% tokens, minutes-long turns). Many capability signals are vendor self-claims and must be read by source.
Update (2026-07-01) — access restored, but only for Fable 5. The June suspension below has been lifted. The US Department of Commerce lifted the export controls on 2026-06-30, and from July 1 Anthropic reopened Fable 5 to users globally on the Claude Platform, Claude.ai, Claude Code, and Claude Cowork (Anthropic — Redeploying Claude Fable 5 · CNBC 2026-06-30). Paid plans included Fable 5 for up to 50% of weekly usage limits through July 7; after that it requires usage credits. Mythos 5 is where the two diverge — it was restored only to a set of US organizations following the government's 2026-06-26 approval, and Anthropic is widening that set through Project Glasswing. So: Fable 5 general access is normal; Mythos 5 remains restricted.
Update (2026-06-12) — access fully suspended (since lifted, see above). Three days after launch, the US government issued an export-control directive under national-security authorities requiring Fable 5 and Mythos 5 to be cut off from all foreign nationals (inside or outside the US, including Anthropic's own foreign-national employees) (Anthropic statement, received 2026-06-12 5:21pm ET). To comply, Anthropic disabled both models for all customers (other models unaffected); the stated cause is that the government became aware of a method to jailbreak Fable 5. Anthropic called this a misunderstanding (a narrow, non-universal jailbreak — prompting the model to review code and find flaws, a capability "widely available from other models"). Source caveat — Anthropic's statement is a primary source but a self-interested party in a dispute with the government, so cross-check the "misunderstanding" framing against neutral press (CNBC, Bloomberg). That it broke in 18 days is itself the axis — access to a closed frontier model is a variable held by regulators, not the vendor.
Verification note (2026-06-11, after an adversarial council audit). The benchmark figures below come from third-party aggregation, and one thing failed verification — the coding gaps (SWE-Bench Pro 80.3 vs 58.6, FrontierCode Diamond 29.3 vs 5.7) mix Anthropic's own-scaffold numbers with neutral-harness numbers, so they aren't on the same axis. On a standardized board where every model shares tooling (Scale SEAL-class), GPT-5.x-class leads ~59% and the ordering reverses. Read coding "dominance" as largely a harness artifact, not a capability gap. Also: this piece was authored through Claude Code about an Anthropic model — a structural conflict of interest — so cross-check pro-Fable framing against non-Anthropic primary sources.
At a glance
| Item | Value |
|---|---|
| Model ID | claude-fable-5 (Mythos 5 = same model, classifier-free, Glasswing-only) |
| Context / output | 1M (default and max) / 128K |
| Price | $10 / $50 per MTok (in/out) — 2× Opus 4.8 ($5/$25) |
| Launch | 2026-06-09, GA on API, Bedrock, Vertex, Foundry — suspended 06-12 by government directive, restored globally 07-01 (see updates above) |
| Reasoning | adaptive thinking always on, raw chain-of-thought never returned (summary only) |
Strengths — what it does well
1. Long-horizon autonomy — the biggest strength
The announcement's headline is "works autonomously for longer than any previous Claude." The flagship example — Stripe completing a codebase-wide migration in a day that would have taken a team over two months (vendor self-claim). Large refactors, overnight coding, and multi-step agent loops where no human is in the loop mid-run are where it fits best. Recommended use — effort high/xhigh with the full spec given up front in the first turn.
2. Coding — strong, but the "dominance" is unverified
Coding is a genuine strength. But the commonly cited gap must be read carefully. Aggregation puts it at SWE-Bench Pro 80.3%, ahead of GPT-5.5 (58.6) and Gemini (54.2) — but that 80.3% is an Anthropic own-scaffold number while GPT's is a neutral-harness number, so they aren't the same axis. On a standardized board where every model shares tooling (Scale SEAL-class), GPT-5.x-class leads ~59% and the best Claude run is ~52% — the ordering reverses. A 5× ratio like FrontierCode Diamond 29.3 vs 5.7 is more suspect still (the same GPT-5.5 scores 92.8 on GPQA). The direction (Fable strong at coding) holds, but the size of the gap is a vendor-scaffold artifact. "Highest" on Cognition FrontierCode and Hebbia finance are vendor-reported.
3. Vision — SOTA, vision-only
Finished Pokémon FireRed with a minimal, vision-only harness, no complex helpers (earlier Claude needed them). Strong on vision-dependent work — screenshot, document, and chart understanding, and computer use.
4. Memory use
Good at writing to and pulling back from file-based memory. In Slay the Spire, memory helped 3× more than for Opus 4.8. Strong at leaving notes and using them in long-running, cross-session agents.
5. Low hallucination + reasoning-token efficiency
On the independent AA-Omniscience benchmark (lower is better), the Claude line is 36.18% vs GPT-5.5's 85.53% — more than 2× lower. The best-anchored of the five figures (AA authored the benchmark). Two caveats — AA's own analysis attributes Claude's low hallucination partly to abstaining/refusing more often (not purely superior knowledge), and 36.18% is a generic "Claude line" figure, so confirm it is Fable 5 specifically. Reasoning efficiency is also a strength — in physics research, reaching nearly where GPT-5.5 landed after four days using a third of the reasoning tokens (vendor-reported).
Weaknesses — what it does poorly / downsides
1. Not the sole leader on science reasoning
On GPQA Diamond, Gemini 3.1 Pro leads (~95.5 on the independent vals.ai evaluation). But the commonly repeated "Fable 3rd, behind GPT-5.5" is inaccurate — on vals.ai, Fable 5 and GPT-5.5 are tied for 2nd at ~93.2. So pure science reasoning isn't Fable's top spot, but the basis for saying it loses to GPT-5.5 is weak. "The smartest model" differs by domain.
2. Safety-classifier over-blocking — some work it just can't do
The most practical weakness. Per TechCrunch, the cyber classifier is keyword-based and broadly blocks legitimate security work. IBM X-Force's Valentina Palmiotti — "rejects even innocuous tasks like reading a blog post," Matt Suiche — "anything in the lexical field of 'cybersecurity' triggers it… a secure-code request is misread as cyber work," another researcher — "even asking for a code review" trips it. When blocked, it returns stop_reason: "refusal" and downgrades to Opus 4.8. For security- or life-sciences-adjacent workloads, over-blocking is a daily cost.
3. Price — 2× Opus 4.8
$10/$50 is exactly 2× Opus 4.8 ($5/$25). For short, repetitive tasks the cost-benefit doesn't hold — Opus 4.8 is enough at half the price. And since the new tokenizer uses ~30% more tokens for the same content, the effective cost gap is wider than 2×.
4. Many operational constraints
- Mandatory 30-day retention, no ZDR — orgs that need ZDR for contract/regulatory reasons cannot use it at all (400 error).
- New tokenizer +~30% tokens — re-baseline max_tokens and context budgets.
- No assistant prefill — code that forced output shape via prefill must switch to structured outputs.
- Minutes-long turns — hard tasks can take several minutes per request. Build timeouts, streaming, and progress UX into the design.
- Refusal-handling overhead — you must treat classifier refusal as a normal state and build fallback paths into the integration.
5. Silent-degradation concern
A weakness flagged by the critique blog (jonready) — the model card states safeguards are invisible to the user and work via prompt modification, steering vectors, or PEFT. On AI-development-adjacent work, if performance quietly degrades, you can't tell whether the model is confused or a policy blocked it. The boundary is vague enough that diagnosis is impossible — a trust problem.
When to use it, when not
Good fits
- Large refactors, overnight coding, multi-step autonomous agents (long-horizon strength)
- Deep 1M-context reasoning, vision-dependent work
- Tasks where accuracy matters more than cost (low hallucination)
Where not to use it
- ZDR-required orgs (outright impossible)
- Security research / offensive work (classifier-blocked)
- Short, repetitive tasks (Opus 4.8 is enough at half the price)
- Top-tier pure science reasoning (Gemini leads)
- Frontier bio/chem research that genuinely needs ungated access → that's Mythos 5 / Glasswing, not Fable
Relationship to Mythos 5
Same model, classifiers removed. Open only to approved Project Glasswing customers (some biology researchers). Identical capability, price, and spec. It lacks the over-blocking weakness above, but general access isn't available. Successor to Mythos Preview.
The suspension — is it really jailbreakable, did the government jump the gun?
A follow-on to the update above, based on 6-angle search, 20 sources, adversarial 3-vote verification. Premise — this story is an asymmetric-source one: nearly every substantive claim originates from Anthropic's self-interested statement during a live dispute (the government has published no technical evidence and hasn't rebutted on the record). What follows is read on top of that asymmetry.
Is it really jailbreakable — yes, but only narrowly
- Narrow jailbreaks are real. Independent researchers (e.g. Pliny the Liberator) bypassed it within 48 hours — layering Unicode/homoglyph substitution, long-context framing, narrative fiction, and decomposition-recomposition to elicit exploit code and chemistry/cyber info (SecurityWeek, cybersecuritynews).
- But no one found a universal jailbreak. Across 1,000+ hours of red-teaming (Anthropic plus reported US-government and UK-AISI review), no method broadly defeats the safeguards. SecurityWeek states plainly these are narrow bypasses and no researcher claims a universal one. cybersecuritynews frames the weakness as structural and transferable across models, not unique to Fable — consistent with Anthropic's "not novel" framing, though that consistency isn't verification (you can't prove a negative — a point dev.to's critique makes correctly).
Did the government jump the gun — on the public record, the optics say yes, but it can't be settled
- Speed / scope: launch (6/9) → directive (6/12), three days; received 5:21pm ET, same-day shutdown for all customers. Scope is all foreign nationals (inside/outside the US, including Anthropic's own staff) — broad relative to a narrow technical basis.
- Thin basis: the letter cited only "national-security authorities," no specifics. Anthropic says the government gave verbal evidence only.
- The decisive wrinkle (Axios, verified 3-0): the actual trigger wasn't a government-demonstrated universal bypass but another company claiming it could jailbreak Mythos. The trigger (a third party's Mythos claim) and the cited basis (a narrow technique the government showed against Fable) are partly conflated, and neither is independently verified.
The one piece of evidence cutting against Anthropic
- The ExploitBench benchmark (techloy / digitalapplied, secondary): Fable 5 ~78% vs GPT-5.5 ~34% — a large offensive-exploit capability gap. If accurate, it directly undercuts Anthropic's "the capability is widely available" defense. Two caveats — it measures offensive exploitation whereas the cited jailbreak demonstrated defensive flaw-fixing (different axis), and the sourcing is blog/aggregation, so low-weight.
One-line verdict
On the public record, the government's action looks disproportionate and premature relative to the technical justification (three days, worldwide foreign-national cutoff, verbal evidence, a rival's claim as the trigger). But the government disclosed no technical evidence, so it can't be settled, and if the ExploitBench gap is real the "widely available" defense wobbles. The core fact is the asymmetry — Anthropic's narrative is effectively the only primary source. Don't conclude from one side's account.
Open (worth tracking). ① the actual content of the government's verbal evidence; ② the identity/verification of the "other company" that claimed a Mythos jailbreak; ③ whether the ExploitBench gap is methodologically sound; ④ what statutory authority BIS invoked, and whether access has been restored or litigated.
Synthesis
In one line — top of the frontier on long-horizon autonomy, coding, vision, and low hallucination, but second-tier on science reasoning, with clear costs in over-blocking, price, and operational constraints. First pick for heavy work where accuracy and autonomy beat cost; for short/repetitive work, security-adjacent work, or ZDR needs, Opus 4.8 or another model is better.
Sources
Primary (Anthropic official)
- Anthropic — Claude Fable 5 and Mythos 5 announcement
- Anthropic — statement on suspending Fable 5 / Mythos 5 access (2026-06-12)
- Introducing Claude Fable 5 (official docs)
- Mythos-class data retention policy (support)
- Project Glasswing
Critique / discourse
Suspension coverage (2026-06-12)
- CNBC — Anthropic disables Fable 5 and Mythos 5 to comply with directive
- Bloomberg — US limits foreign access
- NBC News — Anthropic suspends new models after directive
- Axios — a rival's Mythos-jailbreak claim was the trigger (government-side)
Jailbreak / red-team verification
- SecurityWeek — Anthropic disputes the Fable 5 jailbreak
- cybersecuritynews — Fable 5 bypassed within 48 hours (structural weakness)
- dev.to — narrow bypasses worked, but "nothing actually stopped" critique