Google Launches Gemini 3.8 Live and Extended Thinking Models for Voice AI

Why it matters
Google's latest voice models could redefine user interactions with AI, enhancing productivity and engagement across various sectors.
What happened (in 30 seconds)
- On September 15, 2026, Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced voice processing models.
- These models enable near-real-time reasoning and support 97 languages, addressing previous latency issues in voice AI.
- Developers and enterprises can access these models through the Gemini API, with integration options for partner ecosystems.
The context you actually need
- Prior models struggled with latency during complex tasks, disrupting natural dialogue and user experience.
- Gemini 3.8 Live introduces features like asynchronous function calling and automatic language detection, enhancing conversational capabilities.
- Benchmark results show Gemini 3.8 Live Extended Thinking achieved an 82.6 score on the Speech to Speech Quality Index, indicating superior performance.
What's really happening
Google's Gemini 3.8 Live models represent a significant leap in voice AI technology, addressing long-standing challenges in latency and conversational complexity. The introduction of near-real-time reasoning allows for a more seamless interaction between users and AI, making conversations feel more natural and less robotic. This is particularly important as businesses increasingly rely on voice AI for customer service, virtual assistants, and other applications where human-like interaction is crucial.
The models utilize advanced audio-to-audio architectures, which enable simultaneous speech-and-thought processing. This means that while the AI is responding to a user, it can also think ahead and prepare for follow-up questions or tasks, reducing the lag that often frustrates users. The ability to execute third-party tool calls in the background further enhances functionality, allowing for a richer user experience without interrupting the flow of conversation.
Moreover, the support for 97 languages and visual grounding capabilities expands the potential user base significantly. This global reach means that businesses can deploy these models in diverse markets, catering to a wider audience and improving accessibility. The integration with platforms like Vercel and LiveKit also indicates a strategic move to embed these capabilities into existing workflows, making it easier for developers to adopt and implement the technology.
The competitive landscape is also shifting, as Gemini 3.8 Live positions itself against models from OpenAI and xAI. The high benchmark scores not only validate Google's advancements but also set a new standard for voice AI performance. As developers and enterprises begin to explore these models, the implications for market dynamics could be profound, potentially reshaping how voice AI is perceived and utilized across industries.
Who feels it first (and how)
- Developers: Immediate access to advanced tools for building more responsive voice applications.
- Enterprises: Enhanced customer interaction capabilities, leading to improved service delivery and user satisfaction.
- Consumers: Users of voice-activated devices will experience smoother and more intuitive interactions.
- Global markets: Businesses operating in multilingual environments can leverage the language support to reach broader audiences.
What to watch next
- Adoption rates: Monitor how quickly developers and enterprises integrate Gemini 3.8 Live into their systems, as this will indicate market acceptance.
- Benchmark comparisons: Keep an eye on performance metrics against competitors like OpenAI and xAI to gauge ongoing advancements in voice AI.
- User feedback: Pay attention to consumer experiences and satisfaction levels, as these will shape future iterations and improvements.
Gemini 3.8 Live models are available through the Gemini API and related platforms.
Increased competition in the voice AI market as other companies respond to Google's advancements.
The long-term impact on user behavior and preferences in voice interaction.
Frequently Asked Questions
- Why it matters?
- Google's latest voice models could redefine user interactions with AI, enhancing productivity and engagement across various sectors.
- What happened (in 30 seconds)?
- On September 15, 2026, Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced voice processing models. These models enable near-real-time reasoning and support 97 languages, addressing previous latency issues in voice AI. Developers and enterprises can access these models through the Gemini API, with integration options for partner ecosystems.
- What's really happening?
- Google's Gemini 3.8 Live models represent a significant leap in voice AI technology, addressing long-standing challenges in latency and conversational complexity. The introduction of near-real-time reasoning allows for a more seamless interaction between users and AI, making conversations feel more natural and less robotic. This is particularly important as businesses increasingly rely on voice AI for customer service, virtual assistants, and other applications where human-like interaction is cr
- Who feels it first (and how)?
- Developers: Immediate access to advanced tools for building more responsive voice applications. Enterprises: Enhanced customer interaction capabilities, leading to improved service delivery and user satisfaction. Consumers: Users of voice-activated devices will experience smoother and more intuitive interactions. Global markets: Businesses operating in multilingual environments can leverage the language support to reach broader audiences.
- What to watch next?
- Adoption rates: Monitor how quickly developers and enterprises integrate Gemini 3.8 Live into their systems, as this will indicate market acceptance. Benchmark comparisons: Keep an eye on performance metrics against competitors like OpenAI and xAI to gauge ongoing advancements in voice AI. User feedback: Pay attention to consumer experiences and satisfaction levels, as these will shape future iterations and improvements.
AI news with an enterprise and cloud focus.
"Covers AI in the context of data infrastructure, cloud, and enterprise stacks."
— A47 Editor
Google’s new speech model Gemini 3.8 Live supports real-time reasoning
Google LLC has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its latest speech models designed to tackle latency issues in voice-based AI. These models are capable of near-real-time reasoning and simultaneous speech-and-thought proc...
Daily AI news: models, tools, and policy.
"Independent outlet tracking the fast pace of AI."
— A47 Editor
Google launches Gemini 3.8 Live to take on OpenAI's GPT-Live-1 at a fraction of the cost
Google Deepmind has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new audio models that outperform competitors in the Artificial Analysis speech-to-speech leaderboard, priced at $1.38 per hour, significantly lower than OpenAI's ...
Science and technology stories including AI.
"Longstanding science magazine with thoughtful AI coverage."
— A47 Editor
Google DeepMind wants Gemini to power many different robots
Google DeepMind is advancing its efforts in physical artificial intelligence by integrating its Gemini AI system into humanoid robots, testing its capabilities in real-world environments. This initiative aims to explore how well Gemini can operate in...
Scientific research, technology, environment, and society.
"Scientific American is one of the oldest and most authoritative science magazines, known for deep dives into science, technology, and society."
— A47 Editor
Google DeepMind wants Gemini to power many different robots
Google DeepMind is advancing its efforts in physical artificial intelligence by integrating its Gemini AI system into humanoid robots, testing its capabilities in real-world environments. This initiative aims to explore how well Gemini can operate in...