Trending

    Google Launches Gemini 3.8 Live and Extended Thinking Models for Voice AI

    Section editor: ·Moderate3 articles covering this·3 news sources·Updated an hour ago·World
    Share:
    Infographic comparing Google Gemini 3.8 Live voice models with previous versions and competitors.

    Why it matters

    Google's latest voice models could redefine user interactions with AI, enhancing productivity and engagement across various sectors.

    What happened (in 30 seconds)

    • On September 15, 2026, Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced voice processing models.
    • These models enable near-real-time reasoning and support 97 languages, addressing previous latency issues in voice AI.
    • Developers and enterprises can access these models through the Gemini API, with integration options for partner ecosystems.

    The context you actually need

    • Prior models struggled with latency during complex tasks, disrupting natural dialogue and user experience.
    • Gemini 3.8 Live introduces features like asynchronous function calling and automatic language detection, enhancing conversational capabilities.
    • Benchmark results show Gemini 3.8 Live Extended Thinking achieved an 82.6 score on the Speech to Speech Quality Index, indicating superior performance.

    What's really happening

    Google's Gemini 3.8 Live models represent a significant leap in voice AI technology, addressing long-standing challenges in latency and conversational complexity. The introduction of near-real-time reasoning allows for a more seamless interaction between users and AI, making conversations feel more natural and less robotic. This is particularly important as businesses increasingly rely on voice AI for customer service, virtual assistants, and other applications where human-like interaction is crucial.

    The models utilize advanced audio-to-audio architectures, which enable simultaneous speech-and-thought processing. This means that while the AI is responding to a user, it can also think ahead and prepare for follow-up questions or tasks, reducing the lag that often frustrates users. The ability to execute third-party tool calls in the background further enhances functionality, allowing for a richer user experience without interrupting the flow of conversation.

    Moreover, the support for 97 languages and visual grounding capabilities expands the potential user base significantly. This global reach means that businesses can deploy these models in diverse markets, catering to a wider audience and improving accessibility. The integration with platforms like Vercel and LiveKit also indicates a strategic move to embed these capabilities into existing workflows, making it easier for developers to adopt and implement the technology.

    The competitive landscape is also shifting, as Gemini 3.8 Live positions itself against models from OpenAI and xAI. The high benchmark scores not only validate Google's advancements but also set a new standard for voice AI performance. As developers and enterprises begin to explore these models, the implications for market dynamics could be profound, potentially reshaping how voice AI is perceived and utilized across industries.

    Who feels it first (and how)

    • Developers: Immediate access to advanced tools for building more responsive voice applications.
    • Enterprises: Enhanced customer interaction capabilities, leading to improved service delivery and user satisfaction.
    • Consumers: Users of voice-activated devices will experience smoother and more intuitive interactions.
    • Global markets: Businesses operating in multilingual environments can leverage the language support to reach broader audiences.

    What to watch next

    • Adoption rates: Monitor how quickly developers and enterprises integrate Gemini 3.8 Live into their systems, as this will indicate market acceptance.
    • Benchmark comparisons: Keep an eye on performance metrics against competitors like OpenAI and xAI to gauge ongoing advancements in voice AI.
    • User feedback: Pay attention to consumer experiences and satisfaction levels, as these will shape future iterations and improvements.
    Known:

    Gemini 3.8 Live models are available through the Gemini API and related platforms.

    Likely:

    Increased competition in the voice AI market as other companies respond to Google's advancements.

    Unclear:

    The long-term impact on user behavior and preferences in voice interaction.

    Frequently Asked Questions

    Why it matters?
    Google's latest voice models could redefine user interactions with AI, enhancing productivity and engagement across various sectors.
    What happened (in 30 seconds)?
    On September 15, 2026, Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced voice processing models. These models enable near-real-time reasoning and support 97 languages, addressing previous latency issues in voice AI. Developers and enterprises can access these models through the Gemini API, with integration options for partner ecosystems.
    What's really happening?
    Google's Gemini 3.8 Live models represent a significant leap in voice AI technology, addressing long-standing challenges in latency and conversational complexity. The introduction of near-real-time reasoning allows for a more seamless interaction between users and AI, making conversations feel more natural and less robotic. This is particularly important as businesses increasingly rely on voice AI for customer service, virtual assistants, and other applications where human-like interaction is cr
    Who feels it first (and how)?
    Developers: Immediate access to advanced tools for building more responsive voice applications. Enterprises: Enhanced customer interaction capabilities, leading to improved service delivery and user satisfaction. Consumers: Users of voice-activated devices will experience smoother and more intuitive interactions. Global markets: Businesses operating in multilingual environments can leverage the language support to reach broader audiences.
    What to watch next?
    Adoption rates: Monitor how quickly developers and enterprises integrate Gemini 3.8 Live into their systems, as this will indicate market acceptance. Benchmark comparisons: Keep an eye on performance metrics against competitors like OpenAI and xAI to gauge ongoing advancements in voice AI. User feedback: Pay attention to consumer experiences and satisfaction levels, as these will shape future iterations and improvements.
    3 Articles
    SiliconANGLE — AI

    Google’s new speech model Gemini 3.8 Live supports real-time reasoning

    Google LLC has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its latest speech models designed to tackle latency issues in voice-based AI. These models are capable of near-real-time reasoning and simultaneous speech-and-thought proc...

    THE DECODER

    Google launches Gemini 3.8 Live to take on OpenAI's GPT-Live-1 at a fraction of the cost

    Google Deepmind has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new audio models that outperform competitors in the Artificial Analysis speech-to-speech leaderboard, priced at $1.38 per hour, significantly lower than OpenAI's ...

    Scientific American — Global

    Google DeepMind wants Gemini to power many different robots

    Google DeepMind is advancing its efforts in physical artificial intelligence by integrating its Gemini AI system into humanoid robots, testing its capabilities in real-world environments. This initiative aims to explore how well Gemini can operate in...

    Scientific American

    Google DeepMind wants Gemini to power many different robots

    Google DeepMind is advancing its efforts in physical artificial intelligence by integrating its Gemini AI system into humanoid robots, testing its capabilities in real-world environments. This initiative aims to explore how well Gemini can operate in...