The Agentic Era
No. 26 / 31The Agentic EraProduct2025

Gemini 3

Gemini 3 posted expert-level scores across every major reasoning benchmark and shipped straight into the world's most-used search engine.

Overview

Google released Gemini 3 on November 18, 2025. The benchmark results were unusually broad rather than narrow: a score of 1501 Elo on the LMArena leaderboard, the first model to cross 1500 in blind human preference testing; 37.5% on Humanity's Last Exam without any tool use; 91.9% on GPQA Diamond, a set of graduate-level science questions; and a new state of the art of 23.4% on MathArena Apex. Google described it as the strongest model available for multimodal understanding and its most capable agentic system to date.

The distribution was as significant as the capability. Gemini 3 went into Google Search directly, alongside the Gemini app, AI Studio, and Vertex AI. This placed a frontier reasoning model in front of a user base measured in billions on the day of release, a scale of deployment no other lab could match and one that made the model's behavior a matter of broad public consequence rather than developer interest.

Google launched an agentic development platform called Antigravity alongside the model, aimed at the same long-horizon autonomous work that competitors were building toward. The pattern across the industry by late 2025 was consistent: the frontier model and the agent harness around it were being designed and shipped as one product rather than two.

Key Facts

  • 01Released November 18, 2025, and deployed into Google Search, the Gemini app, AI Studio, and Vertex AI.
  • 02First model to pass 1500 Elo on the LMArena leaderboard, scoring 1501 in blind human preference testing.
  • 03Scored 37.5% on Humanity's Last Exam without tool use and 91.9% on GPQA Diamond, a graduate-level science benchmark.
  • 04Set a new state of the art of 23.4% on MathArena Apex.
  • 05Launched alongside Google Antigravity, an agentic development platform built for long-horizon autonomous coding work.
Why It Matters

Gemini 3 demonstrated that frontier capability and mass distribution were no longer separate problems. For most of the modern era of AI, the strongest models reached a few million paying users while the largest consumer surfaces ran older, cheaper systems. Putting a model at this level into Search collapsed that distinction and made the quality of frontier reasoning a question of general public infrastructure rather than of professional tooling.

The benchmark spread mattered for a second reason. Expert-level performance in isolated domains had been achieved before, but Gemini 3 posted results near or above graduate human level across science, mathematics, multimodal understanding, and general preference simultaneously. Benchmarks that had been designed to last years were saturating within months of publication, and the field's ability to measure its own progress was becoming the binding constraint rather than the progress itself.

The People
Google DeepMind
Sources
[1]

Gemini 3: Introducing the latest Gemini AI model from Google

Google · 2025

https://blog.google/products/gemini/gemini-3/

[2]

Gemini 3: News and announcements

Google · 2025

https://blog.google/products-and-platforms/products/gemini/gemini-3-collection/

[3]

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, et al. · 2025

https://arxiv.org/abs/2501.14249