leejk/ jk lee

Google — Gemini Models · 2026.06

Gemini 3.1 Pro

Strong on reasoning/science (ARC-AGI-2, GPQA), multimodal, and 1M context; weak on agentic/coding, overtaken by its own successor, and operational discontinuity

·
#models#google#gemini#gemini-3-1-pro#landscape

What Gemini 3.1 Pro (Google DeepMind, 2026-02-19) does well and does poorly, as strengths and weaknesses, backed by the official announcement and a competitor table pulled directly for primary figures. Strengths — a reasoning/science core with ARC-AGI-2 verified 77.1% (more than double Gemini 3 Pro), third-party GPQA Diamond science-reasoning #1 (~95.5), Deep Think for hard science/research/engineering, a multimodal foundation, and 1M context. Weaknesses — agentic/coding/documents trail GPT-5.5 (per OpenAI's table: Terminal-Bench 2.0 68.5, GDPval 67.3, OfficeQA Pro 18.1), its own successor Gemini 3.5 Flash (2026-05-19) overtakes it on coding/agentic (Terminal-Bench 2.1 76.2, GDPval-AA 1656, MCP Atlas 83.6), the GPQA #1 is a third-party figure not on the official page, and there is operational discontinuity (Gemini CLI discontinued). Not hands-on — read off the official announcements, the competitor table, and the daily archive, split by source.

TL;DR. Gemini 3.1 Pro (Google DeepMind, 2026-02-19) is a reasoning/science core. StrengthsARC-AGI-2 verified 77.1% (more than double Gemini 3 Pro), third-party GPQA Diamond science-reasoning #1 (~95.5), Deep Think for hard science/research/engineering, a multimodal foundation + 1M context. Weaknessesagentic/coding/documents trail GPT-5.5 (a wide gap on OfficeQA Pro at 18.1), its own successor Gemini 3.5 Flash overtakes it on coding/agentic, the GPQA #1 is third-party, not on the official page, and there is operational discontinuity (Gemini CLI discontinued). Its slot is "not the smartest base model" but "science-reasoning #1, ceding agentic to the successor."

Update (2026-08-15) — 3.5 Pro still has not shipped, which extends this piece's shelf life rather than ending it. Google previewed 3.5 Pro at I/O (2026-05-19) as arriving "next month," yet on 2026-07-21 it released three more Gemini models without it — reporting puts it months behind while coding, math, and hallucination reliability are fixed. So the "being overtaken by its successor" framing below is scoped to the Flash tier. The current Pro-tier model is still 3.1 Pro, and the reasoning/science cell this piece is about was never replaced.

Provenance note. This is not a hands-on review of Gemini 3.1 Pro. The Google official announcements (3.1 Pro, 3.5, Omni) were pulled directly via a headless browser. How to read it — the one primary figure the official announcement carries directly is ARC-AGI-2 77.1%. The GPQA Diamond #1 (~95.5) is third-party (vals.ai) and is not on the official 3.1 Pro page. The agentic/coding numbers come from a competitor table — OpenAI's GPT-5.5 announcement listing Gemini 3.1 Pro as a comparison — so read them as a competitor's own-harness figures. No measured or felt observations.

At a glance

Item Value
Model Gemini 3.1 Pro (Gemini 3 series) — core reasoning baseline
Release 2026-02-19 (developer preview, Vertex AI, Gemini Enterprise, Gemini app, NotebookLM)
Positioning "a smarter model for tasks where a simple answer isn't enough" (official)
Reasoning (ARC-AGI-2, verified) 77.1% — more than double Gemini 3 Pro (official headline)
Science reasoning (GPQA Diamond) ~95.5 (1st)third-party (vals.ai), not on the official page
Context 1M (Gemini 3 series)
Generation position being overtaken on agentic/coding by its successor Gemini 3.5 (2026-05-19)

Why this page — an empty slot, and a stale one

The models overview ranked the frontier base models as Fable 5 · GPT-5.5 · Gemini 3.1 Pro, and Fable 5 and GPT-5.5 got standalone pieces. Gemini was named but never written up on the vendor page — and in the meantime the generation turned over. This piece fills the gap while reconciling the Feb-2026 3.1 Pro → May-2026 3.5 transition.

One framing first. Gemini 3.1 Pro's slot is not "strongest everywhere" but one cell: reasoning/science #1. The models table closed it that way (GPQA 1st), and the official announcement positions it that way (ARC-AGI-2). Coding/agentic was never this model's slot, and now it is ceding even that to its own successor.

Strengths — what it does well

1. Reasoning/science — ARC-AGI-2 77.1% is the official headline

The firmest primary strength. The official announcement headlines ARC-AGI-2 verified 77.1% on a benchmark measuring "the ability to solve entirely new logic patterns"more than double Gemini 3 Pro's reasoning performance. Unlike Fable 5 and GPT-5.5 as ranked on this blog, 3.1 Pro is the one base model that headlines reasoning, not coding.

2. Science-reasoning #1 — but a third-party figure

The basis for this blog ranking Gemini as solo #1 on GPQA Diamond (~95.5) in the models table (vals.ai independent eval; Fable and GPT-5.5 tied 2nd at ~93.2). Important caveat — this GPQA #1 is a third-party figure not on the official 3.1 Pro page. The official headline is ARC-AGI-2; GPQA #1 is a separate axis backed by an independent board — don't fold the two into one source.

3. Deep Think — the hard science/research/engineering core

The official announcement frames 3.1 Pro as "the upgraded core intelligence that makes Gemini 3 Deep Think's breakthroughs possible." Deep Think is the mode for hard science/research/engineering problems, with 3.1 Pro as its baseline. Code-based animated SVGs, complex system synthesis, 3D interactive builds — turning reasoning into practical applications — is the recommended slot.

4. Multimodal + 1M context

The Gemini 3 series is natively multimodal by design (official) and carries 1M context. The same long-context + thinking axis already ranked in the coding-agents model table.

Weaknesses — what it does poorly

1. Agentic/coding/documents trail GPT-5.5

The GPT-5.5 official table listed Gemini 3.1 Pro as a comparison, and it trails across agentic/coding/documents — Terminal-Bench 2.0 68.5 (GPT-5.5 82.7), GDPval wins/ties 67.3 (84.9), SWE-Bench Pro Public 54.2 (58.6, though Opus 4.7 leads at 64.3), OfficeQA Pro 18.1 (GPT-5.5 54.1, Opus 43.6 — last by a wide margin), Toolathlon 48.8 (55.6). This is OpenAI's own-harness competitor table, so weigh the gap size accordingly. The direction is clear — 3.1 Pro is a reasoning model, not an agentic leader.

2. Being overtaken on that slot too, by its own successor

The most important timing weakness. With Gemini 3.5 (2026-05-19, I/O) shipping as "frontier intelligence with action," the official announcement states 3.5 Flash directly outperforms Gemini 3.1 Pro on coding/agentic benchmarksTerminal-Bench 2.1 76.2%, GDPval-AA 1656 Elo, MCP Atlas 83.6%, multimodal CharXiv Reasoning 84.2%, plus 4x the output tokens/sec of other frontier models at less than half the cost. In other words, a successor aimed precisely at 3.1 Pro's weaknesses (agentic/coding/speed/cost) already shipped (3.5 Flash out, 3.5 Pro to follow). The reason to newly adopt 3.1 Pro now narrows to the reasoning/science cell.

3. The GPQA #1 isn't official — split the source

The flip side of strength 2. The "Gemini science-reasoning #1" this blog cites most is not a figure the official announcement backs — it's vals.ai third-party. The official headline is ARC-AGI-2 77.1%; GPQA is a separate board. Mistaking the #1 claim for an official primary inflates the sourcing.

4. Operational discontinuity — Gemini CLI discontinued

A real operational-axis cost. The 3.1 Pro announcement listed Gemini CLI as an access path, but that Gemini CLI was discontinued on 2026-06-18 and force-migrated to Antigravity CLI (daily archive). The transition itself drew "bait and switch" criticism (614 points on HN). An adoption path wobbling along with the generation turnover is a closed-base-model operational risk.

Which job, which model

  • Science reasoning / new logic patterns (ARC-AGI-2) / Deep Think → Gemini 3.1 Pro. This cell is still its strength.
  • Agentic / coding / long-horizon autonomy → now the successor Gemini 3.5 Flash (directly beats 3.1 Pro) or GPT-5.5.
  • Enterprise document / knowledge work → GPT-5.5 (OfficeQA Pro 54.1 vs 3.1 Pro 18.1).
  • Speed / cost sensitive → Gemini 3.5 Flash (4x faster, under half the cost) beats 3.1 Pro.
  • Omnimodal generation (video) → the same vendor's Gemini Omni (a separate line).

Conclusion

Gemini 3.1 Pro is a model with one clear cell.

  1. Reasoning/science — ARC-AGI-2 77.1% is the official headline, and the (third-party) GPQA #1 backs this blog table's science-reasoning summit. This cell is real.
  2. Agentic/codingnever this model's slot (trails GPT-5.5), and now its own successor 3.5 Flash overtakes even that.
  3. Sourcing/operations — the #1 claim is third-party, not official, and the adoption path wobbles with the generation turnover (Gemini CLI discontinued).

One-line take — the science-reasoning #1 cell is firm, but agentic/coding/speed/cost have already passed to its own successor (3.5 Flash). Choosing a Google base model splits as reasoning/science → 3.1 Pro, agentic/coding/cost → 3.5 Flash — and with 3.5 Pro still absent as of 2026-08, that split persists in the odd form of tiers competing inside one generation.

Sources

Primary (Google official — body pulled directly)

Competitor table / independent eval

Same topic