TL;DR. GPT-5.5 (
gpt-5.5, OpenAI, 2026-04-23) is OpenAI's base line, shipped as "our smartest and most intuitive to use model yet — a new class of intelligence for real work." Strengths — agentic/terminal coding (Terminal-Bench 2.0 82.7% SOTA), enterprise document and knowledge work (OfficeQA Pro 54.1%, far above Opus 4.7 and Gemini), price efficiency (API $5/$30 — half Fable 5's input, "AA Coding Index SOTA at half the cost"), and science-reasoning signals (FrontierMath; a Ramsey-number proof verified in Lean). Weaknesses — on raw SWE-Bench Pro Public, Opus 4.7 leads (64.3 vs 58.6, with a memorization footnote), the third-party hallucination benchmark is weak at 85.53%, and it holds no solo #1 domain (BrowseComp, MCP Atlas, FinanceAgent go to Opus/Gemini). The official table compares to Opus 4.7 and Gemini 3.1 Pro, not Fable 5. Read the capability signals by source.
Update (2026-08-15) — GPT-5.6 took the flagship slot. The analysis below is a 2026-06 reading. GPT-5.6 shipped on 2026-07-09 (after a limited 06-26 preview — the public rollout was delayed by government restrictions), split into three variants: Luna, Terra, and Sol. The two axes of this piece were absorbed straight into them — price efficiency went to Terra (GPT-5.5-class performance at half the cost) and agentic/terminal coding to Sol (80 on the AA Coding Agent Index v1.1, at under half the output tokens and roughly a third of the cost). So strengths 1 and 3 below still hold, but the subject changed. Conversely, the weakness noted as no solo #1 domain is partly resolved: Sol claims the cybersecurity cell (the GPT-5.6-Cyber derivative landed 08-10).
Provenance note. This is not a hands-on review of GPT-5.5. But the official OpenAI announcement (2026-04-23) was pulled directly via a headless browser to back the benchmark tables, pricing, and context with primary figures. How to read it — the benchmarks below are OpenAI's own self-reported launch numbers, compared against GPT-5.4, Claude Opus 4.7, and Gemini 3.1 Pro (Fable 5 shipped later, 2026-06-09, and is not in this table). Being vendor self-report, neutral-board reconciliation is still needed, and items the announcement omits (GPQA, hallucination) are split out as third-party (vals.ai, AA-Omniscience). No measured or felt observations.
At a glance
| Item | Value |
|---|---|
| Model ID | gpt-5.5 (base) — variants: gpt-5.5 pro, gpt-5.5 codex, gpt-5.5-cyber; prior gen gpt-5.4 |
| Release | 2026-04-23 ChatGPT/Codex, 04-24 API (System Card alongside) |
| Positioning | "our smartest and most intuitive to use model yet — a new class of intelligence for real work" (official) |
| Price (API) | $5 / $30 per MTok (input/output) — gpt-5.5-pro is $30 / $180. Batch/Flex at half, Priority at 2.5x |
| Context | API 1M / Codex 400K. Codex Fast mode: 1.5x faster for 2.5x cost |
| Efficiency | Matches GPT-5.4 per-token latency while using fewer tokens on the same Codex tasks. Co-designed on NVIDIA GB200/GB300 NVL72 |
| Comparison baseline | Official table is vs Opus 4.7 · Gemini 3.1 Pro (not Fable 5) |
Why this page — a slot the index only named
The models overview ranked the frontier base models as Fable 5 · GPT-5.5 · Gemini 3.1 Pro, and Fable 5 got a standalone piece. GPT-5.5 was named but never written up in the overview table and the OpenAI vendor page. This piece fills that gap — this time with the official announcement pulled directly for primary figures.
One fact to set down first. OpenAI's official table compares to Opus 4.7 and Gemini 3.1 Pro — not Fable 5. Fable 5 shipped a month and a half later (2026-06-09). So the "GPT-5.5 vs Fable" numbers this blog cites elsewhere are all third-party cross-comparisons, while GPT-5.5's primary data is against the Opus 4.7 generation. Miss that gap and the comparison drifts.
Strengths — what it does well
1. Agentic / terminal coding — Terminal-Bench 2.0 SOTA
The firmest primary strength. Per the official announcement, Terminal-Bench 2.0 82.7% — well ahead of Opus 4.7 (69.4) and Gemini 3.1 Pro (68.5) — which OpenAI cites as the basis for "our strongest agentic coding model." On Expert-SWE (an internal long-horizon eval with a 20-hour median estimated human completion time) it beats 5.4 at 73.1% (68.5), and OSWorld-Verified 78.7% (Opus 78.0), CyberGym 81.8% (Opus 73.1). The best fit is command-line / long-horizon work that needs planning, iteration, and tool coordination. The Codex ecosystem (Wasmer, Ramp, Nextdoor, Warp) rides on this (wasmer, nextdoor).
2. Enterprise document and knowledge work — dominant on OfficeQA Pro / GDPval
The "one measured edge" this blog flagged via Databricks' OfficeQA Pro record now has a primary number — GPT-5.5 54.1% vs Opus 4.7 43.6 vs Gemini 3.1 Pro 18.1. A wide gap. GDPval (wins or ties) 84.9% is far above Gemini (67.3), and Investment Banking Modeling (internal) 88.5%. Documents, spreadsheets, and business reasoning are where GPT-5.5's measured edge is clearest.
3. Price efficiency — half Fable's input, "SOTA at half the cost"
A direct contrast to Fable 5's price weakness ($10/$50). GPT-5.5 is API $5/$30 — exactly half the input, 60% the output. OpenAI claims "state-of-the-art intelligence on the AA Coding Index at half the cost of competitive frontier coding models." And it holds GPT-5.4's per-token latency while using fewer tokens on the same Codex tasks — the opposite direction from Fable's +30% tokenizer. On cost-sensitive bulk workloads this axis is large.
4. Science-reasoning signals — FrontierMath / a Ramsey proof
There are primary signals in pure reasoning too. FrontierMath Tier 1–3 51.7% (Gemini 36.9), Tier 4 35.4% (Gemini 16.7), ahead of Gemini (though 5.5 Pro is higher on Tier 4 at 39.6). The announcement cites an internal GPT-5.5 with a custom harness finding a new proof of a longstanding asymptotic fact about off-diagonal Ramsey numbers, later verified in Lean — a mathematical argument, not just code or explanation. Leading among published scores on GeneBench/BixBench (genetics/bioinformatics), a "bona fide co-scientist" claim. But solo #1 in science reasoning is still Gemini (see weaknesses).
5. The quiet operational edge — it doesn't carry what Fable carries
Fable 5's weakness list is GPT-5.5's relative advantage. Fable carries 30-day forced retention, no ZDR, +30% tokens from a new tokenizer, documented classifier over-blocking complaints (TechCrunch), and minute-scale latency. GPT-5.5 is lighter on retention, ZDR, tokens, and price. But it is not "classifier-free" — the announcement explicitly says it is "deploying stricter classifiers" (cyber, biology). So the documented over-blocking complaints are Fable's, while GPT-5.5 has tightened its own cyber safeguards too.
6. Per-job variants + enterprise distribution
Base + Pro (harder problems, higher accuracy) + Codex + Cyber (Trusted Access) + prior 5.4, splitting by job. Thick distribution — AWS Bedrock, Gartner 2026 coding-agent leader, Samsung company-wide, the Ona acquisition.
Weaknesses — what it does poorly
1. On raw SWE-Bench Pro Public, it trails Opus 4.7
The "strip the scaffold and GPT leads coding" generalization from elsewhere is benchmark-dependent. In OpenAI's own table, SWE-Bench Pro (Public) goes to Opus 4.7 at 64.3%, ahead of GPT-5.5 (58.6) and Gemini (54.2). So on one clean public coding benchmark GPT-5.5 is not #1. And OpenAI itself footnotes that benchmark with "labs have noted evidence of memorization" — a warning not to read the figure as a capability gap. GPT-5.5's coding strength is on the Terminal-Bench / agentic side, not a SWE-Bench Pro lead.
2. Hallucination — weak by third-party measure
The official announcement does not report a hallucination figure (AA-Omniscience appears only as an AA Index component). By third-party measure, AA-Omniscience 85.53% (lower is better) is more than double the Fable line's 36.18%. Not backed by the primary announcement, but on accuracy-over-cost work you have to build in abstain and cross-checking.
3. No solo #1 domain — tools, finance, browsing go to Opus/Gemini
Even within the official table, GPT-5.5 doesn't sweep every cell. BrowseComp 84.4 is below Gemini (85.9) and its own 5.5 Pro (90.1); MCP Atlas 75.3 is below Opus (79.1) and Gemini (78.2); FinanceAgent v1.1 60.0 is below Opus (64.4) and 5.4 Pro (61.5). Pure science reasoning (GPQA Diamond) solo #1 is Gemini 3.1 Pro. GPT-5.5's slot is "strong on agentic / documents / price and upper-tier across the board," not "champion of one domain" — it jars if you expect "the smartest model."
4. Numbers are vendor self-report — neutral-board reconciliation remains
This piece's primary figures are OpenAI's own launch numbers, compared to Opus 4.7 and Gemini 3.1 Pro. Not a direct comparison to Fable 5, and not a neutral standard board (Scale SEAL class) where every model runs the same harness. Just as Fable 5 is overrated by its self-scaffold, any vendor table mixes in its own favorable harness/prompts (this one even reports only its own Tau2-bench original-prompt 98.0 while omitting other labs evaluated with prompt adjustment). The direction is primary-backed, but hold the size of any gap until a neutral board settles it.
Which job, which model
- Agentic / terminal / computer-use coding → GPT-5.5 (Terminal-Bench 2.0 SOTA). But know that on the raw SWE-Bench Pro Public head-to-head, Opus 4.7 leads.
- Enterprise document / knowledge work → GPT-5.5 (dominant on OfficeQA Pro / GDPval).
- Cost-sensitive bulk inference → GPT-5.5 ($5/$30, half Fable's input + token efficiency).
- Accuracy over cost (low hallucination required) → GPT-5.5 is weak by third-party measure. Fable is better.
- Top-tier pure science reasoning → Gemini 3.1 Pro (solo #1 on GPQA). GPT-5.5 leads FrontierMath but not GPQA.
- Tool orchestration (MCP) / finance agents / browsing → Opus 4.7 / Gemini are above in the official table.
Conclusion
This time the ranking is read off the primary announcement, pulled directly.
- Coding — agentic/terminal is SOTA (Terminal-Bench 82.7), but raw SWE-Bench Pro Public goes to Opus 4.7, and OpenAI itself footnotes that benchmark for memorization. "Coding lead" is only true once you split by domain.
- Price / documents — $5/$30 (half Fable's input) + OfficeQA Pro dominance is GPT-5.5's cleanest primary strength. The practical pick for enterprise and cost-sensitive workloads.
- Sourcing — primary numbers are backed, but vendor self-report with Opus 4.7 as the baseline, so a direct Fable 5 comparison and a neutral-board reconciliation remain.
One-line take — strong on agentic coding, enterprise documents, and price, but not a raw SWE-Bench Pro leader and weak on hallucination. Ask for "the single smartest model" and the answer splits by domain; GPT-5.5 takes the agentic / documents / price axis.
Sources
Primary (OpenAI official — body pulled directly)
- Introducing GPT-5.5 (official announcement / benchmark tables / pricing / context, 2026-04-23, API 04-24)
- GPT-5.5 System Card (primary safety / safeguard figures)
- Databricks brings GPT-5.5 to enterprise agent workflows (OfficeQA Pro SOTA adoption)
- GPT-5.5 with Trusted Access for Cyber (Cyber variant)
- OpenAI frontier models and Codex on AWS
- Gartner 2026 — agentic coding leader
- Samsung Electronics — ChatGPT / Codex company-wide deployment
Aggregates / independent eval (items the primary announcement omits)
- vals.ai — GPQA board (GPQA Diamond — not in the announcement table)