On September 15, 2026, Google introduced two new real-time voice AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Released just one day ago, they represent the company’s most recent Gemini model announcement currently available for review.
The significance of the announcement goes beyond making AI-generated speech sound more natural. The new models are designed to process visual context, use external tools, and carry out complex, multi-step tasks while continuing a real-time conversation with the user.
However, the functions and performance figures presented by Google may vary depending on the product, account, and deployment environment. Not every feature is immediately available to every user. This article examines what has changed and what limitations remain, based on Google’s official announcement, developer documentation, and model card.
Two Models Designed for Different Types of Interaction
Google introduced two versions of its new live audio model.
Gemini 3.8 Live is designed for fast, fluid real-time voice conversations, with an emphasis on scalability and cost efficiency. In addition to listening and responding through speech, the model can process visual information from a camera or other visual input in near real time.
Gemini 3.8 Live Extended Thinking is intended for more complex tasks that require multi-step reasoning. According to Google, the model can continue speaking with the user while handling a task in the background. It may acknowledge a request with a brief verbal response and provide progress updates as the task proceeds.
For example, if a user requests a reservation involving several conditions or asks the AI to create a set of materials, the model can explain what it is currently doing instead of remaining silent until the final result is ready.
This does not mean that the model reveals its complete internal reasoning process. The “live progress narration” described in Google’s announcement is better understood as a way of communicating the status of a task rather than disclosing the model’s private chain of thought.
Running Tools Without Ending the Conversation
One of the most notable features of Gemini 3.8 Live is its ability to execute external tools and API calls in the background while continuing a conversation.
Previous voice AI systems often paused while searching for information or calling an external service. Google says its new model can acknowledge the request, continue interacting with the user, and process the necessary tools or functions in the background.
Through the Gemini Live API, developers can build applications such as shopping assistants, customer service agents, educational tutors, real-time interpreters, and voice interfaces for vehicles or smart glasses. The Live API processes continuous streams of audio, images, and text and returns spoken responses in real time.
According to Google’s developer documentation, the API also supports “barge-in,” which allows the user to interrupt the model while it is speaking. Other documented functions include tool use, function calling, Google Search integration, and transcription of both user input and model output.
The Model Can Process Visual Context as Well as Speech
Gemini 3.8 Live is not limited to audio. According to the official Google DeepMind model card, the models accept audio, images, video, and text as inputs, with a context window of up to 128,000 tokens. They can produce audio and text outputs, with a stated maximum output length of 64,000 tokens.
This allows the model to respond to visual information shown through a camera or displayed on a screen. Google demonstrated possible applications such as real-time employee onboarding, visual troubleshooting, and discussing a chess game while examining the board.
Visual capability, however, does not guarantee that the model will identify every object or situation correctly. Google’s model card explicitly states that Gemini 3.8 Live and Extended Thinking may exhibit the general limitations of foundation models, including hallucinations.
Language Support Requires Closer Examination
Google’s announcement states that Gemini 3.8 Live can automatically detect and transition between 97 supported languages during a conversation.
However, the Gemini Live API documentation, updated on September 15, lists conversational support for 70 languagesand real-time voice translation in more than 70 languages.
The official materials do not clearly explain why these numbers differ. The range of languages recognized by the model may differ from the number officially supported for conversation or translation through the developer API. Language availability may also vary across the Gemini app, Google Workspace, Search Live, and the Live API.
It would therefore be inaccurate to state that every service and deployment environment fully supports all 97 languages. Developers and users should consult the latest documentation for the particular Gemini product they plan to use.
How Should the Performance Figures Be Interpreted?
Google reports that Gemini 3.8 Live Extended Thinking achieved a score of 82.6 on Artificial Analysis’ Speech-to-Speech Quality Index.
The company also reports a task-completion score of 68.6 percent on τ-Voice and 35.1 percent on Sierra’s τ-Voice banking benchmark. On Big Bench Audio, a benchmark for audio-based reasoning, the model reportedly achieved 97.7 percent.
Google states that Gemini 3.8 Live placed second in the Speech Agent Arena, which evaluates user preferences for voice agents.
These figures provide information about the model’s capabilities, but they do not guarantee the same level of performance in every real-world situation. Benchmarks measure different aspects of a model under specific conditions. Google also notes that one of its voice-agent evaluations was conducted using the Live API on the Gemini Enterprise Agent Platform.
The results should therefore not be interpreted as proof that Gemini 3.8 Live is universally the best voice AI system.
SynthID Watermarking for AI-Generated Audio
Google says audio generated by its AI products includes an imperceptible SynthID watermark. SynthID embeds a signal into the audio output to help identify it as AI-generated.
The technology is intended to address risks such as impersonation and misinformation as synthetic speech becomes increasingly natural. However, a watermark does not automatically prevent every form of misuse, nor does it independently verify the authenticity of all audio found online.
It should be viewed as one transparency measure rather than a complete solution to AI-generated audio abuse.
Is It Available to Everyone?
Google states that Gemini 3.8 Live is rolling out to developers through the Gemini API and Google AI Studio. It is also being introduced to general users through Search Live. For enterprise customers, the model is currently available through a private preview in Gemini Enterprise.
Gemini 3.8 Live Extended Thinking is also rolling out through the Gemini API and Google AI Studio. Google says it is available through Gemini Live, while access through Google Workspace depends on the user’s subscription and the individual service.
Google AI Pro and Ultra subscribers are included in the rollout for Docs, while availability in Gmail and Keep applies to Google AI subscribers under the conditions described in the announcement.
The important wording is “rolling out,” not “fully available to everyone.” Access may differ by country, account, subscription, product, and rollout schedule.
Limitations Disclosed in the Official Model Card
The Google DeepMind model card states that Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking may generate false or unsupported information. It also notes that users may encounter occasional latency or timeout issues.
The documented knowledge cutoff is January 2025. Questions about events occurring after that date may require Google Search or another connected external tool. Answers based solely on the model’s internal knowledge should not automatically be treated as current.
Additional review is necessary when real-time conversations involve personal, medical, or financial information. Although Google presents healthcare and financial services as potential Live API use cases, this does not mean the models can replace qualified medical or financial professionals.
The model card also states that Gemini 3.8 Audio is based on Gemini 3 Pro. Google describes Gemini 3.8 Audio as the collective name for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.
Voice AI Is Moving From Answering Questions to Performing Tasks
The most important change in Gemini 3.8 Live is not simply that its voice sounds more natural. The model represents a shift toward AI systems that can receive spoken and visual information, continue a live conversation, and use external tools at the same time.
Extended Thinking, in particular, is designed for situations in which a task cannot be completed immediately. Rather than remaining silent while processing the request, the model can acknowledge the user and provide spoken progress updates as the work continues.
Still, polished product demonstrations should not be treated as evidence that the same performance will appear in every real-world environment. The language-support figures differ across Google’s official materials, access conditions remain complex, and known limitations such as hallucinations, latency, and timeouts have not disappeared.
Based on the information currently available, Gemini 3.8 Live should not be described as a fully autonomous voice assistant capable of solving every problem. A more accurate assessment is that it is Google’s latest model demonstrating how real-time conversation and tool execution are beginning to converge within a single AI system.

