AI

Google launches Gemini 3.8 Live and Extended Thinking voice models

Google's new voice-first models split scale and reasoning workloads, with the Extended Thinking variant claiming the top spot on Artificial Analysis' speech quality index at 82.6.

T
By TechQuire Daily Staff TechQuire Daily Staff
September 16, 2026 / 7 min read

Voice artificial intelligence has moved from a novelty to a core interface for how people interact with software. For years, conversational agents relied on a rigid pipeline: speech recognition, then text reasoning, then speech synthesis. That architecture introduced latency and made natural interruptions difficult. Google has been pushing against those limits with live dialogue models, and its latest release marks a significant step toward voice systems that can think and talk at the same time.

On September 15, 2026, Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of voice-first models that the company calls its most advanced live dialogue models yet. The launch follows Gemini 3.1 Flash Live, which arrived in March and established a baseline for low-latency conversation. The new models split the problem into two tracks: one optimized for scale and cost, the other for complex, multi-step reasoning. Together, they are designed to give developers and enterprises what Google describes as the building blocks for reliable, production-ready voice agents.

The timing reflects a broader shift in the AI industry. Voice is no longer just an accessibility feature or a hands-free convenience. It is becoming the primary way that users search, draft documents, take notes and control workflows. Google's own product lineup shows this: the company recently introduced Gmail Live for conversational search, Docs Live for draft generation and editing, and Keep Live for note creation.

Competition in the voice AI space is intensifying. Artificial Analysis maintains a Speech to Speech Quality Index that ranks models on how well they handle real-time spoken interaction. Separate benchmarks such as tau-Voice and Sierra's tau-Voice-banking measure whether a model can complete agentic tasks, like booking a reservation or resolving a banking request, through voice alone. Google's new models are aimed squarely at those leaderboards, and the company is making aggressive claims about where they land.

Key Facts

The Google Blog reported on September 15, 2026 that Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are the company's most advanced live dialogue models yet. The models were introduced by Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff, writing on behalf of the Gemini Audio Team. Google describes Gemini 3.8 Live as built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, with increased intelligence and multi-step reasoning.

The headline benchmark is the Speech to Speech Quality Index from Artificial Analysis. Google says Gemini 3.8 Live Extended Thinking captured the number one overall spot with a score of 82.6. On agentic task completion, the model leads with 68.6 percent on tau-Voice and 35.1 percent on Sierra's tau-Voice-banking benchmark. It also scored 97.7 percent on Big Bench Audio while maintaining what Google calls a highly competitive price point versus other frontier models. The base Gemini 3.8 Live secured second place in the Speech Agent Arena and is described as highly cost-effective.

Capabilities go beyond raw scores. Gemini 3.8 Live processes visual inputs in near real time, allowing it to answer live questions about what it sees. It automatically detects and transitions between 97 supported languages mid-conversation. It executes tools and API calls in the background while continuing the conversation, so a user does not have to wait in silence. Extended Thinking reasons and speaks simultaneously, using early verbal cues such as 'Let me check that...' and live progress narration for multi-step background tasks.

Developer access comes through the Gemini Live API, which uses stateful WebSocket connections. According to the documentation, the API handles raw 16-bit PCM audio at 16kHz input, JPEG images at up to one frame per second, and text input, and delivers 24kHz audio output. Platforms including LiveKit, Vercel, Pipecat and Agora manage media streaming. Enterprise partners such as Salesforce, Genspark and Lumeris are testing the technology. Google also confirmed that all model-generated audio includes imperceptible SynthID watermarks.

Rollout began on September 15. For developers, both models are in the Gemini API and Google AI Studio. For enterprises, they are in private preview in Gemini Enterprise. Gemini 3.8 Live is available to everyone in Search Live, as confirmed by Rajan Patel, VP of Engineering for Search, who said the model is already live globally in Search Live. Paid Workspace subscribers gain Extended Thinking features across Docs, Gmail and Keep. On ServiceNow's EVA-Bench, the models push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality, run on the Live API on Gemini Enterprise Agent Platform.

Analysis

The voice AI market has been waiting for a model that can reason without making the user wait. Extended Thinking attempts to collapse that trade-off by verbalizing its process. Saying 'Let me check that...' is not just a courtesy; it is a latency management technique that keeps the conversation alive while background function calls complete. Android Headlines reported on September 15, 2026 that Google demos show the model converting raw paper sketches into working React components and managing restaurant reservations via asynchronous function calls. Those examples illustrate the direction: voice is becoming a control layer for real work, not just a chat interface.

What this really means is that Google is trying to own the full stack of conversational AI, from the underlying model to the developer API to the consumer surfaces where people actually talk to it. The benchmarks matter because they give enterprises a way to compare vendors, but the deeper play is distribution. Gemini 3.8 Live is already in Search Live for everyone, and Extended Thinking is reaching Workspace subscribers through Docs, Gmail and Keep. That combination of a free consumer entry point and paid productivity features creates a funnel that competitors will find hard to match.

9to5Google reported on September 15, 2026 that the launch followed Gemini 3.1 Flash Live from March, and that Extended Thinking is meant for high-complexity tasks such as Gmail Live, Docs Live and Keep Live. The report also noted that the model began rolling out to Gemini Live that day. This cadence suggests Google is iterating rapidly, roughly every six months, and using each release to push further into agentic territory. The base model's second place in the Speech Agent Arena is also telling: Google is not just chasing raw quality, it is segmenting the market between cost-sensitive deployments and premium reasoning workloads.

Unite.AI reported on September 15, 2026 that the models are rolling out across the Gemini API, Google AI Studio, Gemini Enterprise, Search Live, Gemini Live and Google Workspace. The breadth of that list is unusual for a single launch. It signals that Google has unified its voice roadmap across consumer, developer and enterprise teams, rather than shipping separate models for each. The bigger picture here is that voice is becoming a platform layer, and the companies that control the default voice assistant on phones, browsers and productivity suites will have a durable advantage in the next wave of AI adoption.

Why It Matters

Voice interfaces are finally crossing the threshold from demos to daily use. The ability to switch between 97 languages mid-conversation, answer questions about live visual input, and run background tasks without breaking the flow changes what a voice assistant can be. For developers, it means a single API that handles audio, vision and tool use over a stateful connection.

Google's emphasis on SynthID watermarking is also significant. As synthetic voice becomes indistinguishable from human speech, provenance and detection become critical for trust. By watermarking all generated audio, Google is trying to get ahead of misuse concerns while still shipping powerful capabilities. That balance will shape regulation and user acceptance in the months ahead.

The competitive landscape is watching closely. A number one spot at 82.6 gives Google a marketing advantage, but rivals will respond with their own updates. The real test is whether the models hold up in production, where latency, cost and reliability matter more than a single benchmark run.

Next Up

Both models are rolling out now. Developers can access them through the Gemini API and Google AI Studio, while enterprises can join a private preview in Gemini Enterprise. Gemini 3.8 Live is available to everyone in Search Live, and Extended Thinking is reaching Gemini Live and Docs for Google AI Pro and Ultra subscribers. Paid Workspace subscribers are getting Extended Thinking features across Docs, Gmail and Keep.

Expect the next phase to focus on deeper agentic workflows. Google has already demonstrated sketch-to-React conversion and asynchronous reservation booking. The roadmap likely includes more third-party integrations through partners such as Salesforce, Genspark and Lumeris, and further optimization of the Live API for low-latency, multimodal interactions. As voice becomes the default interface for AI, the race will shift from benchmark scores to everyday reliability at scale.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.