Models·2 min read
By BitsMindsSource: Google DeepMind

Google Gemini 3.1 Ultra Breaks Reasoning Records with 94.3% on GPQA Diamond

Google's flagship Gemini 3.1 Ultra achieves a verified 94.3% on GPQA Diamond and 77.1% on ARC-AGI-2, setting new state-of-the-art benchmarks for complex scientific and logical reasoning.

Google Gemini 3.1 Ultra Breaks Reasoning Records with 94.3% on GPQA Diamond
Share:

Google released Gemini 3.1 Ultra in March 2026 as the flagship of its new three-model family, alongside Flash-Lite and Flash Live. While the broader Gemini 3.1 family had been anticipated for its multimodal improvements, it is the Ultra's reasoning benchmark performance that has drawn the most attention from the AI research community: a verified 94.3% on GPQA Diamond — a benchmark designed to stump even PhD-level domain experts — and 77.1% on ARC-AGI-2, which evaluates a model's ability to solve entirely novel logical patterns it has never encountered before.

The 3.1 Ultra improves on its predecessor in several key dimensions. It retains the 2-million-token context window that made Gemini 3 a landmark release, but adds significantly improved grounding capabilities to reduce hallucinations on factual queries. A new sandboxed Code Execution tool allows the model to write, run, and debug code mid-conversation without external plugins, bringing an integrated software engineering loop directly into the chat interface. Deep Think mode — available to Google AI Ultra subscribers — enables extended deliberative reasoning for problems that require sustained multi-step analysis.

The model operates natively across text, image, audio, and video without requiring transcription intermediaries, a capability Google calls Multimodal Mastery. This means the model processes raw audio and video frames directly, enabling more accurate comprehension of tone, visual context, and temporal relationships than systems that first convert media to text before reasoning. In early enterprise tests, this has proven especially valuable for medical imaging analysis, video content moderation, and complex document processing that mixes charts, diagrams, and prose.

Gemini 3.1 Pro is rolling out broadly to Gemini app users with AI Pro and Ultra subscriptions, and is available on NotebookLM. Developers and enterprises can access both Pro and Ultra in preview through the Gemini API via AI Studio, Vertex AI, and Gemini CLI. With Gemini now at 750 million monthly users, the 3.1 family represents Google's most aggressive push yet to close the gap with OpenAI in both consumer mindshare and enterprise deployments — a competition that has become one of the defining technology races of 2026.

More on Gemini

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Many kinds of input, one Qwen context On a deep violet field, a filmstrip, an audio waveform, a landscape photograph and a text page curve toward a single luminous sphere bearing the official purple Qwen star. A label below reads “1M context”, representing text, images, audio and video sharing one million-token context window. Aa 1M context BITSMINDS.COM
Models

Qwen3.8-Omni-Flash Undercuts Gemini on Audio and Video

Hill Climb: GPT-6 Astra versus Claude Fable 5.1 A teal desert buggy climbs a sandstone ridge on the left. An orange jeep climbs a green hill on the right, under an arc of gold coins. The two landscapes meet at a diagonal divide. 7 GPT-6 ASTRA CLAUDE FABLE 5.1 VS HILL CLIMB BITSMINDS.COM
Models

GPT-6 Astra vs Claude Fable 5.1: Hill Climb

Gemini 3.8 Live: thinking while the conversation continues An editorial illustration in a dark blue and violet room. A carefully drawn studio microphone and a tilted smartphone flank a translucent speech bubble carrying the multicoloured Gemini star. A continuous luminous audio waveform travels between them. Above the bubble, a separate arc connects small search, reasoning and completion symbols, representing background work continuing during a spoken conversation. The phone screen and visual paths are conceptual, not a reproduction of Google's actual interface or internal reasoning. Original vector illustration for BitsMinds, gemini-3-8-live-extended-thinking-voice. 16 September 2026. GEMINI 3.8 LIVE EXTENDED THINKING Gemini 3.8 LIVE The conversation continues Thinking. Still talking. BITSMINDS.COM
Models

Gemini 3.8 Live Tops Voice AI and Undercuts GPT