What Is Gemini 3.5 Transcribe? Google Speech to Text Guide

2026-08-27
Google's Gemini 3.5 Transcribe brings low-latency speech to text across 85+ languages with filler word removal and multi-speaker support. Here's what it does.
Google DeepMind released Gemini 3.5 Transcribe on August 26, 2026. It's their most accurate speech to text model to date, and it's already running on products millions of people use daily. If you've ever dictated a message and watched it come out garbled, this is the model built to fix that problem.
What Is Gemini 3.5 Transcribe?
Gemini 3.5 Transcribe is a speech to text model built on Google's Gemini 3.5 architecture. It converts spoken audio into written text in real time or after recording. The model handles 85+ languages with automatic detection, so you don't have to manually pick one before talking.
The difference between this and older Google speech to text models goes beyond accuracy. This one cleans up what you say. It strips filler words like "um" and "ah" from the final output. It handles self-corrections too. If you say "let's meet Tuesday, no, Wednesday," the transcript reads "let's meet Wednesday." That's a big deal for anyone who dictates messages or records meetings.
According to Artificial Analysis benchmarks, the model hits a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode. Lower is better. These numbers represent a meaningful jump over Google's previous Chirp 3 model, especially in latency-sensitive scenarios like live voice typing where every millisecond of delay is noticeable.
How Google's AI Transcription Model Works
Under the hood, Gemini 3.5 Transcribe processes audio in chunks rather than waiting for you to finish speaking. That's what streaming means here. The model listens, transcribes, and refines as you talk.
The pipeline does several things in sequence. First, it identifies the language being spoken. Then it converts audio frames into text tokens. After that, a cleanup pass strips filler words, fixes self-corrections, and applies formatting like punctuation and capitalization. All of this happens in the time between you finishing a sentence and the text appearing on screen.
One part that's genuinely useful for professionals: custom vocabulary. You can feed the model specialized jargon, product names, or unique spellings. It'll recognize those terms instead of guessing at phonetically similar words. If you work in a technical field where standard speech to text models butcher your terminology, this fixes that problem directly.
There's also multi-speaker identification. The model can distinguish between up to three speakers and label their turns with timestamps. More than three speakers moves into experimental territory. Don't rely on it for panel discussions just yet.
Key Features of the Gemini Transcription Model
85+ language support with automatic detection. The model figures out what language you're speaking without manual configuration. Useful if you switch between languages mid-conversation.
Custom vocabulary recognition. Provide a word list and the model transcribes those terms correctly. Medical professionals, developers, and legal workers benefit most from this.
Filler word removal. "Ums" and "ahs" get stripped from the final text automatically. You talk naturally, the output reads clean.
Self-correction handling. When you correct yourself mid-sentence, the model keeps the final version and drops the false start. No more "I meant" artifacts cluttering your transcripts.
Function calling. This is where things get interesting. The model can trigger other Gemini models to act on what you said. Mention "generate an image of a cat" and the transcription model can pass that instruction to an image generation model. It transcribes and delegates. Two jobs in one pass.
Multi-speaker labeling. Up to three speakers get identified with timestamps. Beyond three is experimental.
Where Gemini 3.5 Transcribe Already Runs
This isn't a lab demo. Google has already shipped it to several products you can use right now.
Gboard Rambler on Android. If you use Gboard, you've got access to voice typing that auto-formats your speech. The Rambler feature takes your raw dictation and produces clean, formatted text. This is the most direct way to try Gemini 3.5 Transcribe on your phone. You just talk, and the model handles the rest.
Gemini app for macOS. The desktop Gemini app has a "Speak to Window" capability. You talk, it transcribes and responds. The integration shows how the model works alongside other Gemini features rather than functioning as a standalone transcription tool.
Google Antigravity. The prompt box microphone in Google's Antigravity platform uses this model for voice input. Developers can try it through Google AI Studio, which offers API access to the model in public preview.
Chrome browser (coming soon). Google says Gemini 3.5 Transcribe will power talk-to-type in any text field across Chrome. That means dictating into web forms, search boxes, and comment sections without installing anything extra. If you spend time filling out forms online, this could replace typing for routine tasks.
If you want to try the voice features on mobile, the Google Gemini Android app is your best entry point. It ties together the broader Gemini family of capabilities, including AI transcription powered by this model.
Gemini 3.5 Transcribe vs Chirp 3
Google's previous speech to text workhorse was Chirp 3, part of the Cloud Speech-to-Text API. Gemini 3.5 Transcribe isn't a minor update to that model. It's a different architecture built on the Gemini 3.5 foundation.
The headline number: 70% improvement in final transcription latency compared to Chirp 3. You wait less between finishing your sentence and seeing the clean output. For real-time use cases like voice typing, that's the difference between usable and frustrating.
On the FLEURS benchmark, which tests cross-lingual speech recognition across dozens of languages, Gemini 3.5 Transcribe scores 5.50% WER streaming and 5.04% WER non-streaming. These are competitive numbers for a model handling 85+ languages. Chirp 3 was solid for English but dropped off significantly in less-represented languages.
Custom vocabulary is new territory. Chirp 3 didn't have anything comparable. If you needed domain-specific terms recognized with Chirp 3, you'd post-process the transcript or build custom language models. Gemini 3.5 Transcribe handles this natively, which saves time and effort for anyone working with specialized terminology.
Function calling is another addition. Chirp 3 transcribed audio and handed off the text. The new model can delegate tasks to other Gemini models based on what it heard. That's a workflow shift, not just an accuracy bump.
Who Benefits From This AI Transcription Model
Casual users. If you dictate texts, emails, or notes, you'll notice cleaner output with less manual editing. The filler word removal and self-correction handling do most of the heavy lifting here. You talk the way you normally would, and the model produces text you can send without revising.
Professionals who need accurate transcripts. Journalists, researchers, and legal workers who deal with interviews or recorded meetings get a model that handles multiple speakers and recognizes specialized terminology. The custom vocabulary feature matters most for this group. Instead of manually fixing every misrecognized technical term, you set up your vocabulary list once and the model handles the rest.
Developers. The model is available in public preview through the Gemini API via Google AI Studio and Google Antigravity. You can build it into your own apps. The function calling capability opens up possibilities for voice-driven interfaces that go beyond simple dictation. If you're building an app that needs voice input, this is worth testing.
Enterprise users. Google is bringing this to the Gemini Enterprise Agent Platform and Gemini Enterprise for CX (customer experience). If your company uses Google's enterprise tools, expect to see Gemini 3.5 Transcribe powering voice features in those workflows soon.
What to Watch Out For
Multi-speaker identification is reliable for up to three speakers. Beyond that, it's experimental. If you need to transcribe a four-person roundtable, you'll get results, but accuracy drops. Plan accordingly.
The model auto-formats text. That's great until it isn't. If you need a verbatim transcript that preserves every "um" and false start for research or legal documentation, this model will fight you on that. There's no indication of a verbatim mode yet, so check your output carefully if precision matters.
Custom vocabulary requires setup. You can't just start talking and expect it to know your company's internal acronyms. You need to provide the word list beforehand through the API or app settings. Once configured, it works well, but it's not automatic.
The 85+ language support is impressive, but WER varies across languages. English will perform better than lower-resource languages. Google hasn't published per-language benchmarks. If accuracy in a specific language matters to your workflow, test it before committing. Privacy-conscious users should also review how audio data is processed and stored when using the API or cloud-based transcription features.
The Takeaway
Gemini 3.5 Transcribe fixes the things people actually complained about with Google speech to text: latency, filler words, self-corrections, and multi-speaker confusion. It's already running in Gboard and the Gemini app, so you can try it today. Download the Google Gemini app from APKPure to explore the voice features powered by this model, and see if AI transcription finally works the way you need it to.