Models·4 min read
By BitsMindsSource: Google

Gemini 3.8 Live Tops Voice AI and Undercuts GPT

Google’s new speech-to-speech pair takes the top of the Artificial Analysis quality index — and the cheaper sibling runs at about a seventh of GPT-Live-1 Astra’s hourly cost. Both reason in the background without stopping the conversation.

Gemini 3.8 Live: thinking while the conversation continues An editorial illustration in a dark blue and violet room. A carefully drawn studio microphone and a tilted smartphone flank a translucent speech bubble carrying the multicoloured Gemini star. A continuous luminous audio waveform travels between them. Above the bubble, a separate arc connects small search, reasoning and completion symbols, representing background work continuing during a spoken conversation. The phone screen and visual paths are conceptual, not a reproduction of Google's actual interface or internal reasoning. Original vector illustration for BitsMinds, gemini-3-8-live-extended-thinking-voice. 16 September 2026. GEMINI 3.8 LIVE EXTENDED THINKING Gemini 3.8 LIVE The conversation continues Thinking. Still talking. BITSMINDS.COM
Share:

Google has released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of native speech-to-speech models that take the top spot on Artificial Analysis’ Speech to Speech Quality Index while costing less per hour than the OpenAI and xAI models they beat. Both went live on 15 September through the Gemini API and Google AI Studio, with a consumer rollout landing the same day in Search Live, the Gemini app and Workspace.

The split between the two is a deliberate one. Gemini 3.8 Live is the volume model — built for scale and cost, and aimed at the kind of always-on voice interface that runs for hours. Extended Thinking is the one that reasons while it talks: it can work through a multi-step problem in the background and keep speaking to you at the same time, filling the gap with verbal acknowledgements rather than the dead air that gives current voice assistants away.

Where the numbers land

Voice quality and agentic voice workHigher is better · ET = Extended Thinking · Artificial Analysis and Sierra, 15 Sept 2026 · n/r = not reportedGemini 3.8 Live ETGPT-Live-1 AstraGrok Voice Think Fast 2.002040608010082.681.581.3Speech to SpeechQuality Index68.667.956.5tau-Voice(agentic)35.132.0n/rtau-Voicebanking (Sierra)
Google takes the top of the speech-to-speech table by 1.1 points over GPT-Live-1 Astra, and by 2.7 points on Sierra’s banking suite. The margins are real but narrow; the standard Gemini 3.8 Live model, not charted here, scores 76.0 on the quality index. Data: Artificial Analysis Speech to Speech Quality Index.

Extended Thinking scores 82.6 on the Speech to Speech Quality Index, ahead of GPT-Live-1 Astra on 81.5 and Grok Voice Think Fast 2.0 on 81.3. On agentic voice work the lead widens a little: 68.6% on tau-Voice and 35.1% on Sierra’s banking variant, against 67.9% and 32.0% for Astra. Google also reports 97.7% on Big Bench Audio, an audio reasoning assessment.

These are close margins on a leaderboard that has changed hands repeatedly this year, and it is worth saying plainly that a 1.1-point quality lead is not the story. The prices are.

The cost line is the actual news

What an hour of talking costsUS$ per hour of audio input — lower is better · Artificial Analysis, 15 Sept 2026Gemini 3.8 LiveGemini 3.8 Live ETGrok Voice TF 2.0GPT-Live-1 Astra01234560.83.54.85.8Cost per hour ofaudio
The price gap is wider than the quality gap. The model that wins the leaderboard runs at about 60% of GPT-Live-1 Astra’s hourly cost; the cheap sibling that scores 76.0 runs at about a seventh of it. Data: Artificial Analysis.

Google lists the 3.8 Live models at $0.005 per minute of audio input and $0.018 per minute of audio output. Measured on Artificial Analysis’ cost-per-hour basis, that puts Extended Thinking at $3.50 an hour against $5.83 for GPT-Live-1 Astra and $4.80 for Grok Voice Think Fast 2.0. The cheaper Gemini 3.8 Live runs at $0.84 an hour — roughly a seventh of Astra — for a quality index of 76.0.

That is the trade a buyer actually has to make, and it is not the one the headline implies. If you want the best voice model available, it is now Google’s and it costs 40% less than the previous best. If you want a voice model that is merely good enough for a support line, Google will sell you one at a price that makes the comparison awkward for everyone else.

What the models can do that the last generation could not

Three capabilities are new rather than incremental. The models detect and switch language mid-conversation across 97 languages, holding accent consistency rather than resetting to a default voice. They execute tool and API calls asynchronously in the background, so a lookup no longer stops the dialogue. And they take near real-time visual input, which is what makes Search Live work — you point a camera at something and keep talking about it.

The DeepMind model card puts the pair on a Gemini 3 Pro base with a 128K-token context window, 64K tokens of output, and a January 2025 knowledge cutoff. All generated audio carries a SynthID watermark. A third model shipped alongside them: Gemini 3.5 Transcribe, a speech-to-text model covering 85-plus languages at a 4.0% word error rate streaming and 2.6% non-streaming.

Where it shows up

Developers get both models in the Gemini API and AI Studio, with launch integrations for LiveKit, Pipecat, Agora, LangChain and Vercel. Enterprises get Extended Thinking in a private preview inside Gemini Enterprise, with Salesforce, ServiceNow and Lumeris named as early deployments. Consumers get it in Search Live for everyone, and in Gemini Live plus Gmail, Docs and Keep for AI Pro and Ultra subscribers.

Google has been pushing at this slot all year, from Gemini 3.5 Live Translate in June to a summer spent trailing OpenAI’s GPT-Live, which launched in July. It also now has a distribution channel it did not have then: the Gemini-powered Siri that shipped with iOS 27 on 14 September. A cheaper, faster voice model underneath that assistant is a compounding advantage, and it is the part of this launch that will still matter in six months.

More on Gemini

Evergreen coverage we keep current — start here.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Atria Dawn Preview 744B · MIT · shipped with no announcement BITSMINDS.COM
Models

Atria Dawn: A 744B Agent Model, Shipped in Silence

Anthropic model cadence: waiting for the next beat A sculpted terracotta metronome with the Claude symbol stands on a warm cream surface. Beside it, a solid model cartridge reads Fable 5.1, released 1 September. An outlined, translucent future cartridge reads Fable 5.2 with a prominent question mark and the word Unconfirmed. Six small solid beats represent the six models released since April. The image illustrates a release pattern and an unconfirmed rumour, not an announced model or launch date. ANTHROPIC MODEL CADENCE / 2026 RELEASE RHYTHM ANTHROPIC The next beat? A release pattern. A rumour. An open question. FABLE 5.1 RELEASED / 01 SEP 2026 FABLE 5.2 ? UNCONFIRMED NO OFFICIAL ANNOUNCEMENT SIX RELEASES SINCE APRIL BITSMINDS.COM
Models

No Fable 5.2 Yet, but Anthropic’s Cadence Says Soon

SAKANA AI · FUGU MAX + ULTRA V2 One API, a pool it won’t name BITSMINDS.COM
Models

Sakana's Fugu Ultra v2 Routes Around Astra and Fable