Google has introduced two new models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, aimed at improving near-real-time reasoning, voice-agent capabilities and natural interactions with AI.
The new models are designed to enable developers and enterprises to build production-ready voice agents capable of handling conversational interactions while performing complex tasks in the background. The models also bring more fluid voice interactions to Gemini across the Gemini app, Google Workspace and Search.
Gemini 3.8 Live is designed for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking, meanwhile, targets more complex tasks by combining conversational capabilities with enhanced multi-step reasoning.
According to Google, Gemini 3.8 Live Extended Thinking secured the top overall position on Artificial Analysis' Speech to Speech Quality Index, with a score of 82.6. It also recorded 68.6 per cent on the τ-Voice agentic task-completion benchmark and 35.1 per cent on Sierra's τ-Voice-banking benchmark. The model scored 97.7 per cent on Big Bench Audio, highlighting its capabilities in audio-based reasoning.
Gemini 3.8 Live is positioned as a more cost-efficient option for developers and enterprises seeking to deploy voice agents at scale. The model also ranked second in the Speech Agent Arena, according to Artificial Analysis.
Beyond voice interaction, Gemini 3.8 Live can process visual inputs in near real time, allowing conversations to incorporate visual context. It can automatically detect and switch between 97 supported languages during conversations, while also executing tools and API calls in the background.
The ability to continue conversations while tasks are being processed is designed to make voice-based AI interactions more fluid and collaborative, particularly for applications requiring multiple steps or external tool use.
With the two models, Google is expanding Gemini's capabilities from conversational AI towards more real-time, multimodal and agentic interactions, providing developers with infrastructure for voice agents that can reason, interact with visual information and perform tasks while maintaining an ongoing conversation.

