Gemini 4: When It Ships and Why Google Skipped 3.5 Pro
Gemini 4 finished pre-training and is now in post-training, and DeepMind's new chief says it will ship "much earlier" than the end of 2026. Where the model stands, why Gemini 3.5 Pro never launched, what went wrong at Google this summer, and how far Gemini 4 has to climb to reach the top of the leaderboard.
Gemini 4 has finished pre-training and is now in post-training, the stage where a raw model is tuned for instruction-following, tool use and safety before anyone outside Google gets it. Koray Kavukcuoglu, who took over day-to-day leadership of Google DeepMind in August, said so on stage at The Information’s AI Agenda Live summit this week, and gave the first real signal on timing: Google hopes to ship it “much earlier” than the end of the year. “Our intention is to … release an early post-training output because we see the results and we are excited,” he said, according to 9to5Google. There is still no date, no model ID, no pricing and no benchmark table.
Where Gemini 4 stands today
The confirmed facts are few, and they come in two dates. On 21 July, at the bottom of the post announcing Gemini 3.6 Flash, Google wrote that it had “already started our most ambitious pre-training run yet, for Gemini 4.” A day later Sundar Pichai repeated the line on the Q2 earnings call. Two months on, Kavukcuoglu says that run is done. Google is already testing the model internally in Antigravity, its agentic coding environment, and safety testing is under way, The Decoder reports. Pichai has said what the model is for: Google needs “much larger base models” to compete, and coding and agentic coding are where it has the most catching up to do. He framed the target as competing “at the frontier level of where the frontier will be when Gemini 4 comes out,” Search Engine Journal reported.
When will it be released?
Google has not committed to a date, so any window is an estimate. Ours starts from three things. Kavukcuoglu spoke of an “early post-training output”, which sounds like a preview rather than a finished, generally available model; Google has launched its last several flagships that way, with Gemini 3 Pro arriving as a preview in November 2025. Pre-training took about two months to reach post-training, and the company is under obvious pressure to move fast. And “much earlier than the end of the year”, said in late September, leaves roughly October and November. Our reading is a Gemini 4 preview in the Gemini app and AI Studio somewhere between late October and November, with general availability and cheaper Flash variants following over the next few months. Kavukcuoglu promised “fast-paced iterations” after the first release, which fits that pattern. Treat the specific window as inference, not reporting.
Did Google really skip Gemini 3.5?
Not the whole generation, just its flagship. Gemini 3.5 Flash launched at I/O on 19 May, and 3.5 Flash-Lite and a security-tuned 3.5 Flash Cyber followed in July. What never shipped is Gemini 3.5 Pro, which Pichai announced at the same I/O for a June release. The version numbers then kept climbing through the Flash line: 3.6 Flash in July, 3.7 Flash in August, and 3.8 Flash on 2 September, alongside the 3.8 Live voice models. The result is an odd line-up in which Google’s newest models are all mid-tier, and its newest Pro-tier model is still Gemini 3.1 Pro from the spring.
Asked directly about 3.5 Pro, Kavukcuoglu said Google “took a little bit of a step back” to focus on Flash models, and that the priority is now Gemini 4. He did not say it was cancelled. In July, TechCrunch reported that Logan Kilpatrick still expected 3.5 Pro to “land soon”, and Pichai told investors it was “currently in testing”. Neither has happened since. In practice the model has been leapfrogged: the compute, the people and the attention have moved to the next generation.
What went wrong with 3.5 Pro
Google has never given a technical reason, so what is known comes from reporting. According to Bloomberg, cited by TechCrunch, the model “struggled to meet internal performance goals.” More detailed accounts, summarised by Tech Times from HackerNoon’s unnamed sources, describe two failures in an earlier near-finished build: it broke down in recursive tool-calling, where an agent calls one tool and uses the result to call the next, and it could not keep complex, multi-layered SVG layouts structurally consistent. Google reportedly decided fine-tuning could not fix the first problem and restarted training. A later delay was blamed on frequent hallucinations and inconsistent performance in real workflows. The public deadlines went one after another: June, then 30 June, then 17 July. We covered the slips as they happened, from the first miss to the third.
Recursive tool-calling is not a niche skill. It is the core of agentic coding, the market where Anthropic and OpenAI have pulled furthest ahead this year, and it is exactly the area Pichai later singled out as the one Gemini 4 has to fix.
A rough summer for Google
The model was not Google’s only problem. In June it lost Gemini co-lead Noam Shazeer to OpenAI and AlphaFold Nobel laureate John Jumper to Anthropic within three days. In July it was rationing Meta’s Gemini capacity as compute ran short. On 5 August, Axios reported, Demis Hassabis stepped aside as DeepMind’s chief executive to become its chairman and Alphabet’s chief scientist, with Kavukcuoglu taking over operations and reporting to Pichai; the same day, Jeff Dean left to found a start-up. The Decoder, citing the Guardian, reports that Gemini development is now being consolidated in the Bay Area. And this month Google opened Claude Opus 5 to all its engineers, a telling choice for a company whose own flagship was supposed to be the best coding model it could offer.
Kavukcuoglu’s answer to all of it was confidence. “In my mind, it’s a certainty that we are always gonna be at the frontier,” he said. He also moved away from his predecessor’s language about artificial general intelligence, telling the audience the real question is whether Google can “build intelligent agents that we can trust”.
How good does Gemini 4 need to be?
The distance is easy to measure. On Artificial Analysis' Intelligence Index, which averages a set of reasoning, knowledge, coding and agentic tests, Google’s best model today is Gemini 3.8 Flash at 41. Claude Opus 5.5, released on 22 September, leads at 58, with Claude Fable 5.1 and GPT-6 Astra tied on 53. Gemini 3.1 Pro, the last Pro model Google actually shipped, sits at 30.
To be competitive, Gemini 4 has to land in the mid-50s; to lead, it needs to clear Opus 5.5 at the point it ships, which by then may be a moving target. The Flash results suggest Google is not starting from nothing: 3.8 Flash scores 54.9% on HLE-Verified and, by Google’s account, beats most larger frontier models on the DeepSWE engineering benchmark at a fraction of the price. A much larger model trained on that recipe could close a lot of the gap. Expect Google to lead with coding and long-running agent benchmarks, since that is where Pichai says the work has gone, and to keep the 1M-token-plus context windows that have been a Gemini strength. Pricing will matter too: the Flash line has been undercutting everyone, and Google has room to do the same at the top.
What you should not rely on are the numbers already circulating. Since mid-September, posts on X have claimed that a model labelled “gemini-3.8-flash” in Arena’s anonymous battles was really a Gemini 4 Pro checkpoint codenamed “Argon”, with a leaked chart showing 88% on DeepSWE and 86.8% on OSWorld-2.0. An audit of those claims found that the chart has no source or methodology, Arena’s scored leaderboard lists no Gemini 4, and neither Google nor Arena has commented. They may turn out to be close. For now they are unverified.
What to watch
- A preview label. An “experimental” or “preview” Gemini 4 in AI Studio would be the early post-training release Kavukcuoglu described.
- Independent scores. Google’s own launch benchmarks will favour Google. The first Artificial Analysis and Arena results will show whether the gap to Opus 5.5 and GPT-6 has closed.
- Tool calling at length. The failure that sank 3.5 Pro is the one to test first: long, chained agent runs that do not fall over halfway.
- The fate of 3.5 Pro. If Gemini 4 ships and 3.5 Pro is never mentioned again, the question of whether it was cancelled answers itself.
Google has been here before. Gemini 3 arrived as a preview in November 2025 and went to the top of the leaderboards within days. The difference this time is that Google has fewer of the people who built that model, and it is running out of time to deliver on “much earlier than the end of the year”.
More on Gemini
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.