Google Unveils Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking in Major Advancement for Real-Time AI Voice Agents

Google has officially introduced its most advanced live dialogue models to date: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Announced by Principal Engineer Tom Ouyang and Member of Technical Staff Malini Jaganathan on behalf of the Gemini Audio Team, this release marks a significant leap forward in near real-time reasoning, parallel processing, and voice agent capabilities. The new models are engineered to make human-AI collaboration more intuitive, fluid, and capable of executing complex multi-step tasks using natural speech.

The rollout spans consumer applications—including the Gemini app, Google Workspace, and Google Search—as well as robust enterprise tools via the Gemini Live API. By bridging the gap between conversational fluidity and heavy computational reasoning, Google aims to redefine how developers, enterprises, and everyday consumers interact with voice-activated artificial intelligence.
Background Context and Technical Evolution

The launch of the Gemini 3.8 Live series builds upon years of rapid iteration in speech-to-speech architectures. Traditional voice assistants have historically relied on a multi-step pipeline: converting speech to text, processing the text through a large language model, and then synthesizing the text back into speech. While functional, this pipeline often introduced noticeable latency, stripped away emotional nuance, and struggled with maintaining context during rapid interruptions.
Over the past two years, the AI industry has shifted heavily toward native speech-to-speech modalities. Google’s introduction of the Gemini 3.8 Live models addresses the remaining hurdles in this domain: lowering latency, enabling simultaneous thinking and speaking, integrating visual inputs on the fly, and supporting deep asynchronous task execution. By embedding advanced reasoning directly into the audio pipeline, these models can handle sophisticated workflows—ranging from customer service banking benchmarks to writing code—entirely through spoken conversation.

Performance Benchmarks and Empirical Data
Independent evaluations underscore the technical dominance of the new lineup. According to Artificial Analysis, Gemini 3.8 Live Extended Thinking captures the number one overall position on the Speech to Speech Quality Index with a score of 82.6. Furthermore, it demonstrates unprecedented strength in agentic task completion, scoring 68.6% on the Ï₂-Voice benchmark and 35.1% on Sierra’s Ï₂-Voice-banking benchmark. Its reasoning capabilities are further highlighted by a 97.7% score on Big Bench Audio, outperforming many competing frontier models while maintaining a highly competitive cost structure.

Meanwhile, Gemini 3.8 Live has secured the second-place ranking in the competitive Speech Agent Arena. Developers and enterprise clients evaluating these models on ServiceNow’s EVA-Bench—a rigorous testing framework for voice agents—will find that the models successfully push the Pareto Frontier, striking an optimal balance between conversational quality and execution accuracy for complex workflows.
Key Architectural Innovations and Multimodal Capabilities

The Gemini 3.8 Live architecture introduces several critical features designed to remove friction from live interactions:
Near Real-Time Visual Processing: The model can process visual inputs instantaneously. Whether an employee is undergoing real-time onboarding or a user is playing a game of chess against the AI, Gemini 3.8 Live uses visual context to answer questions live while maintaining a natural conversational flow.

Dynamic Multilingual Support: Gemini 3.8 Live automatically detects and seamlessly transitions between 97 supported languages mid-conversation, removing the need for manual language selection.
Asynchronous Tool Execution: The model can execute background tools and API calls while continuing to converse with the user. It acknowledges requests immediately and keeps the dialogue moving forward while tasks finish processing in the background.

Simultaneous Reasoning and Speech: Gemini 3.8 Live Extended Thinking reasons and speaks at the same time. For complex tasks, it uses early verbal cues—such as “Let me check that…”—to naturally acknowledge prompts, while providing live progress narration to walk users through multi-step background operations.
Integration Across the Google Ecosystem

Consumers will encounter these advancements natively embedded across Google’s core product ecosystem. In Google Workspace, features like Docs Live, Gmail Live, and Keep Live allow users to draft documents, manage email threads, and sort through notes using conversational voice commands. Inside Search Live, users can receive step-by-step, real-time troubleshooting help. Additionally, the Gemini app leverages the new models to deliver Daily Briefs, assist in achieving inbox zero, and manage complex daily schedules through verbal hand-offs.
Ecosystem Partnerships and Industry Reactions

To maximize the reach of these models, Google has partnered with a wide array of developer platforms and enterprise infrastructure providers. Through the Gemini Live API, platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents now allow developers to build and deploy high-performance voice interfaces. These platforms manage the complex real-time media streaming infrastructure behind the scenes, enabling developers to focus purely on application logic and user experience.
Major enterprise players have also expressed strong enthusiasm for the release. Representatives from Salesforce, Genspark, Lumeris, 11Sight, Equal AI, ServiceNow, Lenskart, Ambr AI, Casuu, and LiveKit have highlighted the models’ low latency, conversational fluidity, and advanced tool-calling capabilities as key drivers for upgrading their enterprise solutions.

Commitment to Transparency and Safety
In alignment with Google’s ongoing commitment to responsible AI development, all audio generated by Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking is embedded with SynthID watermarking. This imperceptible watermark is woven directly into the audio output, ensuring that AI-generated content remains easily detectable to help mitigate the spread of misinformation. Comprehensive details regarding safety measures, evaluations, and ethical considerations are outlined in the official Gemini 3.8 Audio model card.

Broader Implications and Future Outlook
The commercial rollout of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking signals a maturing market for voice-first artificial intelligence. As enterprises migrate from experimental chatbots to production-ready voice agents, the ability to execute asynchronous tasks, handle complex multi-step workflows, and process visual data simultaneously becomes a vital competitive advantage.

By lowering the barrier to entry for developers through robust API integrations and maintaining aggressive pricing structures, Google is positioning its Gemini ecosystem as a foundational layer for the next generation of voice-driven software. As these models roll out to consumers and enterprises worldwide, they promise to fundamentally alter how humans command, collaborate with, and rely upon artificial intelligence in daily professional and personal workflows.







