Cleaner transcripts arrive via Gemini 3.5 Transcribe’s built-in filler word removal
Professional meetings and lectures often yield messy transcripts filled with “um,” “uh,” and hesitant pauses that clutter the final document. Fixing these errors manually wastes twenty minutes, turning a time-saving tool into a chore. Google tackles this friction with Gemini 3.5 Transcribe, a model designed for high-fidelity audio processing.
The advance extends past basic conversion to include intelligent cleaning. Standard engines aim for literal verbatim output, whereas this model separates meaningful content from verbal filler. It edits audio in real-time, discarding the linguistic debris that makes human speech inefficient for records.
Differences from standard STT
Most Speech-to-Text (STT) engines struggle outside controlled environments. Background noise in coffee shops or offices often triggers hallucinations or causes the model to lose the conversation thread. The Gemini 3.5 family, including Live and Transcribe variants, handles these interruptions.
- Noise Robustness: Precision remains intact even when speech is drowned by ambient sound.
- Multilingual Support: Coverage spans over 85 languages, expanding global AI workflows.
- Jargon Detection: Specialized fields like medicine, law, and technology often confuse models with niche terms; this update improves detection for such jargon.
- Contextual Intelligence: The 3.5 architecture understands sentence flow, enabling filler removal without breaking grammar.
Deployment and Model Variants
Google separates this capability into distinct models for developers integrating them into specific workflows:
- Gemini 3.5 Transcribe: The engine converting messy audio into clean text. It serves users needing documentation from meetings or interviews.
- Gemini 3.5 Live: Optimized for low-latency conversations. It supports fluid, voice-controlled AI where speed matters as much as accuracy.
- Gemini 3.5 Live Experimental: A sandbox for testing real-time audio reasoning limits.
Voice-dependent LLM agents benefit significantly from this shift. Prompt engineers have spent effort writing system instructions to ignore fillers, but handling this at the transcription layer is more efficient. It prevents “garbage in, garbage out” before text reaches the primary LLM.
A Google employee shared an internal document outlining these developments, though it reflects only their opinion, not the whole firm. SemiAnalysis acted as a vessel to share these points, which raise interesting questions about the current state of AI. As an ad-free reader-supported publication, SemiAnalysis highlights that major open problems are already solved for users. For example, people run foundation models on a Pixel 6 at 5 tokens / sec. Meanwhile, OpenAI faces stiff competition as a third faction gains ground in this quiet arms race.
The full Gemini 3.5 Pro release teased earlier this year remains pending. These specialized tools offer practical applications of the architecture, solving daily productivity issues. Users relying on voice memos or recorded meetings should monitor the broader rollout closely.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
I'm curious if Gemini 3 filters filler words live or just scrubs the transcript after it's done. The model essentially edits the audio in real-time, discarding the linguistic debris that makes human speech inefficient for documentation.
Summarizing action items via prompt is a game changer. Does it ever miss the smaller details in the transcript? To get the best results, you should use a model that distinguishes between meaningful content and verbal filler to ensure the output isn't cluttered with linguistic debris.

The cleanup feature is a lifesaver for client calls. Have you tried enabling the "Discard Filler Words" option? This feature essentially edits the audio in real-time, discarding the linguistic debris that makes human speech so inefficient for documentation. Which specific setting makes it work best for you?