What Is Gemini 3.8 Live? Google's New Voice Models Explained

2026-09-16
Google's Gemini 3.8 Live and Extended Thinking models bring real-time speech, live vision, and background reasoning to Android. Here's how the Gemini Live voice agent works.
Google opened September with a voice update that landed across the apps millions of Android owners already have on their home screen. On September 15, the Google DeepMind Gemini 3.8 team introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two models built for talking with an AI the way you talk with a person. Neither one is a chatbot you type into. Both are native speech-to-speech models, which means they hear you, think, and answer out loud without stopping to convert your words into text first.
If you have ever asked your phone a question and waited through an awkward pause while it "thought," this release is aimed at you. The pitch is simple: conversation that keeps moving while the machine does its work in the background.
What Gemini 3.8 Live actually is
Gemini 3.8 Live is a voice model designed for scale and cost. Google describes it as combining conversational intelligence with fluid dialogue and visual grounding. In plain terms, it can hold a real back-and-forth, and it can see what your camera sees while you talk.
The second model, Gemini 3.8 Live Extended Thinking, is the heavyweight. It runs multi-step reasoning while it speaks, so it can grind through a complicated request without going quiet on you.
Both sit inside the Gemini Audio family, and both are based on Gemini 3 Pro. Inputs cover audio, images, video, and text, with a token context window up to 128K and an output window of 64K. That context matters more than it sounds. A long conversation about a document, a recipe, or a repair job stays in memory instead of falling off the edge.
There is no open-weights download here. These are hosted models, so you reach them through Google's servers rather than running them on your phone. That also shapes the privacy picture: your audio goes to Google to be processed, not to a model sitting in local storage.
How the Gemini voice model handles a live conversation
The old way of building a voice assistant was a relay race. One system listened and wrote down what you said. A language model read the transcript and wrote a reply. A third system read that reply out loud. The older architecture chained three separate systems together, and every handoff between them added delay, so a small transcription mistake early in the chain could snowball into a badly wrong answer by the time the voice came back out of the speaker.
Gemini 3.8 Live skips the relay. It hears audio and produces audio in one motion, which is why the conversation feels closer to a phone call than a walkie-talkie.
The practical payoff shows up in three places. First, it processes what your camera points at in near real time, so it can talk about an object it's looking at. Second, it detects and switches between 97 supported languages mid-conversation, and it keeps your accent consistent when it speaks back. Third, it runs tools and API calls in the background, so it can say "give me a second" and keep chatting while it finishes.
That last behavior is the real shift. The model acknowledges your request, keeps the conversation alive, and closes the loop when the job is done. No dead air.
What Gemini 3.8 Live Extended Thinking does differently
Extended Thinking reasons and speaks at the same time. That sounds like a small tweak. It isn't. Older reasoning models either answer fast and shallow or go silent to think hard.
The model uses early verbal cues to hold the floor. It says something like "let me check that" while it works, then narrates progress step by step as a multi-step task runs. For anything that takes more than a few seconds, that chatter is the difference between trusting the system and abandoning the conversation.
Google's own demos show it turning a hand sketch plus spoken feedback into a working React component, and coordinating a multi-step booking. Those are developer-facing examples, but they hint at the shape of things for regular users: you describe what you want, keep talking to refine it, and the model handles the middle.
The benchmark numbers back up the ambition. Extended Thinking took the number one overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. It hit 68.6% on the τ-Voice agentic task benchmark, 35.1% on Sierra's τ-Voice-banking, and 97.7% on Big Bench Audio. The base Live model landed second in the Speech Agent Arena, a human preference test.
Where you can use Gemini 3.8 Live on Android today
The rollout started the day of the announcement, and the reach is wider than a developer preview.
For everyday users, Gemini 3.8 Live powers Search Live in the Google app. Ask a question out loud and the answer comes back as a spoken conversation rather than a page of links. Extended Thinking is rolling out to Gemini Live, and for Google AI Pro and Ultra subscribers it extends into Workspace with Docs, plus Gmail and Keep for Google AI subscribers.
That spread is worth pausing on. The same voice model that runs a general chat also handles conversational search in Gmail and note creation in Keep. If you pay for a Google AI plan, you get smarter voice tools in the apps you already open every morning.
Developers reach both models through the Gemini Live API and Google AI Studio, priced at $0.005 per minute for audio input and $0.018 per minute for audio output.
Practical Android examples of a Gemini Live voice agent
Abstract capability is easy to wave at. Here is what it looks like in your hand.
Picture pointing your camera at a confusing appliance error code while asking what it means. The model reads the display, explains the fault, and walks you through a fix in one conversation. No screenshot, no typing.
Now imagine a kitchen with your hands covered in dough. You ask for a substitution, the model answers, you ask a follow-up, and it remembers the recipe you mentioned two minutes ago. The 128K context is what keeps that thread intact.
Or take a noisy walk in the park. You want directions to a landmark and a quick fact about it. The voice agent answers in one turn, and the background tool calls fetch the map data without cutting the conversation.
The common thread is hands-free, eyes-free use. That's where a real speech-to-speech model beats a text interface, and it's why Google is pushing these models toward phones rather than keyboards.
Benefits and limits of Gemini 3.8 Live
What you gain:
- Natural pacing that survives interruptions and barge-in
- Real-time vision, so the model can discuss what your camera shows
- Automatic language switching across 97 languages in one session
- Background tool execution that keeps the conversation flowing
- Available in Search Live and Gemini Live for regular users, not just developers
What to watch for:
- The model card lists hallucinations and occasional slowness or timeouts as known limits
- The knowledge cutoff date is January 2025, so very recent events may need a live tool call
- Audio output is watermarked with SynthID, so AI-generated clips are detectable
- These are hosted models with no self-hosted option
The hallucination point isn't a footnote. A voice assistant sounds more authoritative than a text box, and that confidence can make a wrong answer harder to spot. Ask for a source when a fact matters. So what should you check before trusting a spoken answer? Anything with a number, a date, or a name attached.
How Gemini 3.8 Live fits next to other voice assistants
Most voice assistants on Android today are either text models bolted onto a speech layer or narrow command parsers. You can tell which one you're talking to within a sentence. The text-model approach stumbles on turn-taking and accent consistency. The command parser falls apart the moment you ask something open-ended.
Gemini 3.8 Live sits between those poles and aims past both. It handles open-ended conversation, keeps a stable voice, and runs real tools mid-sentence. The benchmark scores suggest a real jump. Benchmarks and lived experience don't always agree, so I'd treat the leaderboard numbers as a signal, not a guarantee.
For comparison shopping, the closest rivals are other frontier speech-to-speech systems. The differentiator here is background reasoning that doesn't interrupt the talk, plus a rollout that puts the model inside Search and Gmail rather than a standalone app.
The takeaway on Gemini 3.8 Live
Gemini 3.8 Live is Google's bet that talking to AI should feel like talking to a person, not waiting on a machine. The base model keeps conversation fluid and cheap enough to scale. Extended Thinking adds the muscle for tasks that need real reasoning, and it does that without muting the conversation.
If you use the Google app on Android, you're already partway there through Search Live. If you subscribe to a Google AI plan, the same voice engine reaches Gmail, Docs, and Keep.
The smart move is to try a task you'd normally type, and ask for it out loud instead. Start with something small, like a question about a photo or a quick fact, then push toward a multi-step request and watch how the model narrates its progress. That back-and-forth is where these voice models earn their keep, and it's the clearest way to judge whether the upgrade matches the hype. Download the latest Gemini app from APKPure to get the updated voice experience, and keep an eye on your in-app settings to see which features have switched on for your account.