Google releases Gemini 3
Gemini 3 launches directly in the Gemini app and Google Search, with record scores on several reasoning benchmarks at release.

What happened
On 18 November 2025 Google released its new generation of models, Gemini 3, beginning with Gemini 3 Pro in preview. “This is the first time we are shipping Gemini in Search on day one,” wrote Sundar Pichai, chief executive of Google and Alphabet.
On launch day it reached:
- Everyone in the Gemini app;
- Google AI Pro and Ultra subscribers in AI Mode in Search;
- Developers through the Gemini API in AI Studio, Gemini CLI and the new Google Antigravity coding platform;
- Businesses through Vertex AI and Gemini Enterprise.
TechCrunch noted that it came just seven months after Gemini 2.5, less than a week after OpenAI released GPT-5.1 and about two months after Anthropic released Claude Sonnet 4.5, “a reminder of the blistering pace of frontier model development”.
Background: the third generation
Gemini is the family of models built by Google DeepMind, and also the name of Google’s AI assistant app. In his note, Pichai traced the history: the first Gemini’s breakthroughs were “native multimodality” (understanding text, images, audio and video together) and long context; the second laid the foundations for “agentic” abilities, with Gemini 2.5 Pro topping the LMArena leaderboard for more than six months.
At launch, Google said the Gemini app had more than 650 million monthly users, AI Overviews in Search had 2 billion users a month, and 13 million developers had built with its generative models.
How good is it?
Google published a comparison (see its evaluation methodology) with Gemini 2.5 Pro, Anthropic’s Claude Sonnet 4.5 and OpenAI’s GPT-5.1. The biggest gaps were on two hard reasoning tests.
Gemini scores measured by Google. Gemini 2.5 Pro and Claude Sonnet 4.5 scores are from the Scale AI leaderboard; GPT-5.1 from Artificial Analysis. Source: Google DeepMind, Gemini 3 Pro evaluation methodology and results, Nov 2025
The Deep Think score allowed code execution. Source: Google, “A new era of intelligence with Gemini 3” and evaluation methodology, 18 Nov 2025
Some other results Google published:
| Test | What it measures | Gemini 3 Pro | Comparison |
|---|---|---|---|
| GPQA Diamond | Graduate-level science questions | 91.9% | GPT-5.1: 88.1% |
| MathArena Apex | Very hard competition maths | 23.4% | Claude Sonnet 4.5: 1.6% |
| Terminal-Bench 2.0 | Completing tasks in a command line | 54.2% | GPT-5.1: 47.6% |
| SWE-bench Verified | Fixing issues in real software projects | 76.2% | Claude Sonnet 4.5: 77.2% |
| Vending-Bench 2 | Running a simulated vending-machine business for a year | $5,478 | Claude Sonnet 4.5: $3,839 |
Note the second-to-last row: on SWE-bench Verified, a widely used coding test, Gemini 3 Pro scored slightly below Claude Sonnet 4.5 (77.2%) and GPT-5.1 (76.3%).
TechCrunch reported that Gemini 3’s Humanity’s Last Exam result was the highest on record, beating the previous best of 31.64% held by GPT-5 Pro. (TechCrunch gave Gemini 3’s score as 37.4%, slightly different from the 37.5% in Google’s announcement.)
What it can do
Google’s announcement gave examples, all of them Google’s own demonstrations:
- Learning. Reading and translating handwritten family recipes in different languages into a shareable cookbook; turning academic papers or long video lectures into interactive flashcards and visualisations; analysing a video of your pickleball match and producing a training plan.
- “Generative interfaces” in Search. In AI Mode, Gemini 3 builds visual layouts, interactive tools and simulations on the fly in response to a question, not just a block of text.
- Coding. Google called it the best “vibe coding” model it had built (vibe coding means describing what you want in everyday language and letting the AI write the program). Demos included a retro 3D spaceship game, a playable sci-fi world and interactive web pages.
- Getting things done. Google AI Ultra subscribers could try “Gemini Agent” in the Gemini app for multi-step tasks such as organising a Gmail inbox or booking local services.
Safety
Google called Gemini 3 its “most secure model yet”, saying it had undergone the most comprehensive safety evaluations of any Google AI model. It reported less sycophancy (telling users what they want to hear), more resistance to “prompt injection” attacks and better protection against misuse for cyberattacks. Beyond its own testing, Google gave early access to bodies such as the UK AI Security Institute and had independent assessments from firms including Apollo, Vaultis and Dreadnode.
The stronger Deep Think mode was not released on day one. Google said it was taking extra time for safety evaluations and feedback from safety testers before offering it to Ultra subscribers in the following weeks.
Things to keep in mind
- The results come from Google. Gemini’s scores were measured by Google; rivals’ scores mostly come from their own reports or third-party leaderboards, so test settings differ. Google’s methodology document lists the source for each.
- It didn’t lead everywhere. On SWE-bench Verified it scored slightly below Claude Sonnet 4.5 and GPT-5.1.
- It was a preview. Google released Gemini 3 Pro in preview and kept updating it.
- This model is now almost a year old. Google has since released Gemini 3 Flash (December 2025, which became the Gemini app’s default) and Gemini 3.1 Pro (February 2026, scoring 77.1% on ARC-AGI-2). A planned Gemini 3.5 Pro was eventually cancelled, and the next flagship was Gemini 4 Argon in September 2026.
What it means for ordinary users
- Available in the Gemini app. From launch day, everyone could use Gemini 3 in the Gemini app. For plan prices, see our Gemini product page.
- Search changed too. Google AI Pro and Ultra subscribers got Gemini 3 answers and interactive layouts in AI Mode in Search.
- For coders. It is available in AI Studio, Gemini CLI and Antigravity, and in third-party tools such as Cursor, GitHub, JetBrains and Replit.
Model summary
- Developer
- Google DeepMind
- Availability
- Gemini app, Google Search AI Mode, API
Reported benchmark results
| Humanity's Last Exam (no tools) | 37.5% | Deep Think mode: 41.0% |
| GPQA Diamond | 91.9% | |
| LMArena | 1501 Elo | |
| ARC-AGI-2 | 45.1% | Deep Think mode, with code execution |
Data source: Google announcement. Scores are as published by the developer at launch. Test conditions differ between companies, so compare with care.