Models & Research

Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

· September 15, 2026
Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents

What changed

Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue AI models to date. These models can execute background tasks like API calls and tool operations without pausing the conversation flow. They also handle live visual input and can switch fluently between 97 languages mid-dialogue. Extended Thinking boasts top scores on critical benchmarks, ranking first on the Artificial Analysis Speech to Speech Quality Index with a score of 82.6 and achieving 97.7% on Big Bench Audio. Both models are accessible now via the Gemini API and Google AI Studio at $0.005 per thousand tokens.

Why builders should care

These models represent a significant upgrade for developers building voice agents and live conversational AI. The ability to run tools and APIs seamlessly in the background means voice assistants can perform complex operations without disrupting user interaction or response speed. Handling live visuals adds a new dimension for multimodal applications, such as remote support or real-time recognition tasks. Multi-language switching mid-conversation expands global usability and reduces friction in multilingual environments. Benchmark-leading quality scores highlight a step forward in natural, reliable voice interactions, crucial for production-grade deployment.

The practical takeaway

For teams building voice-powered customer service bots, virtual assistants, or interactive agents, Gemini 3.8 Live models lower technical complexity by integrating real-time API execution and multi-language support in one package. At a competitive $0.005 per thousand tokens, the pricing keeps live agent deployments cost-effective at scale. The extended thinking capability also means AI can keep track of longer, more complex conversations with less user friction. This can translate directly into better user satisfaction, reduced fallback to human agents, and streamlined workflows for voice and multimodal applications.

What to watch next

Observe how Google integrates these models into broader AI and cloud services, potentially bundling with edge compute or enterprise tooling. Track adoption patterns among businesses with global multilingual needs or those requiring real-time visual input processing. It will be critical to see if competitors respond with similar seamless background execution and multimodal strengths or cut pricing to defend market share. Builders should monitor SDK support, latency claims, and real-world integration stories to validate these announced capabilities beyond benchmarks.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.