Gemini 3.8 Live Tops Voice AI and Undercuts GPT
Google’s new speech-to-speech pair takes the top of the Artificial Analysis quality index — and the cheaper sibling runs at about a seventh of GPT-Live-1 Astra’s hourly cost. Both reason in the background without stopping the conversation.
Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of native speech-to-speech models that take the top spot on Artificial Analysis’ Speech to Speech Quality Index while costing less per hour than the OpenAI and xAI models they beat. Both went live on 15 September through the Gemini API and Google AI Studio, with a consumer rollout landing the same day in Search Live, the Gemini app and Workspace.
The split between the two is a deliberate one. Gemini 3.8 Live is the volume model — built for scale and cost, and aimed at the kind of always-on voice interface that runs for hours. Extended Thinking is the one that reasons while it talks: it can work through a multi-step problem in the background and keep speaking to you at the same time, filling the gap with verbal acknowledgements rather than the dead air that gives current voice assistants away.
Where the numbers land
Extended Thinking scores 82.6 on the Speech to Speech Quality Index, ahead of GPT-Live-1 Astra on 81.5 and Grok Voice Think Fast 2.0 on 81.3. On agentic voice work the lead widens a little: 68.6% on tau-Voice and 35.1% on Sierra’s banking variant, against 67.9% and 32.0% for Astra. Google also reports 97.7% on Big Bench Audio, an audio reasoning assessment.
These are close margins on a leaderboard that has changed hands repeatedly this year, and it is worth saying plainly that a 1.1-point quality lead is not the story. The prices are.
The cost line is the actual news
Google lists the 3.8 Live models at $0.005 per minute of audio input and $0.018 per minute of audio output. Measured on Artificial Analysis’ cost-per-hour basis, that puts Extended Thinking at $3.50 an hour against $5.83 for GPT-Live-1 Astra and $4.80 for Grok Voice Think Fast 2.0. The cheaper Gemini 3.8 Live runs at $0.84 an hour — roughly a seventh of Astra — for a quality index of 76.0.
That is the trade a buyer actually has to make, and it is not the one the headline implies. If you want the best voice model available, it is now Google’s and it costs 40% less than the previous best. If you want a voice model that is merely good enough for a support line, Google will sell you one at a price that makes the comparison awkward for everyone else.
What the models can do that the last generation could not
Three capabilities are new rather than incremental. The models detect and switch language mid-conversation across 97 languages, holding accent consistency rather than resetting to a default voice. They execute tool and API calls asynchronously in the background, so a lookup no longer stops the dialogue. And they take near real-time visual input, which is what makes Search Live work — you point a camera at something and keep talking about it.
The DeepMind model card puts the pair on a Gemini 3 Pro base with a 128K-token context window, 64K tokens of output, and a January 2025 knowledge cutoff. All generated audio carries a SynthID watermark. A third model shipped alongside them: Gemini 3.5 Transcribe, a speech-to-text model covering 85-plus languages at a 4.0% word error rate streaming and 2.6% non-streaming.
Where it shows up
Developers get both models in the Gemini API and AI Studio, with launch integrations for LiveKit, Pipecat, Agora, LangChain and Vercel. Enterprises get Extended Thinking in a private preview inside Gemini Enterprise, with Salesforce, ServiceNow and Lumeris named as early deployments. Consumers get it in Search Live for everyone, and in Gemini Live plus Gmail, Docs and Keep for AI Pro and Ultra subscribers.
Google has been pushing at this slot all year, from Gemini 3.5 Live Translate in June to a summer spent trailing OpenAI’s GPT-Live, which launched in July. It also now has a distribution channel it did not have then: the Gemini-powered Siri that shipped with iOS 27 on 14 September. A cheaper, faster voice model underneath that assistant is a compounding advantage, and it is the part of this launch that will still matter in six months.
More on Gemini
Evergreen coverage we keep current — start here.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.