Google brings cleaner AI transcription to Gemini Audio

Google has updated Gemini Audio with Gemini 3.5 Live, 3.5 Live Experimental, and the new Gemini 3.5 Transcribe model. The update focuses on cleaner transcription, stronger language handling, jargon recognition, and better voice-controlled AI features.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

This is a routine product update for transcription quality and voice workflows, with only mild dependence or automation implications.

Google brings cleaner AI transcription to Gemini Audio

Google is expanding Gemini Audio with new Gemini 3.5 models built for speech, dictation, and transcription. The update adds Gemini 3.5 Live, 3.5 Live Experimental, and a new transcription model called Gemini 3.5 Transcribe.

The headline feature is cleaner text from spoken audio. Google says the new tools can recognize specialized jargon, detect more than 85 languages, deal with background noise, and keep working when speech is interrupted.

What Gemini 3.5 Transcribe Adds

Gemini 3.5 Transcribe is a new member of the Gemini family. Its job is to turn speech into text with less cleanup after the fact, especially in situations where people speak naturally rather than in polished, scripted sentences.

Google says that 3.5 Transcribe "represents a major advancement from our previous transcription model, Chirp 3," with particular gains in multilingual performance and wording error rates.

The model can automatically format text and remove filler words like "um" and "uh." That matters because many transcripts are useful only after someone edits them into readable form. By handling those basic cleanup steps, Gemini 3.5 Transcribe is aimed at making the first draft closer to something people can use.

Google also says users can "edit naturally with just your voice." In practical terms, that positions transcription as more than a passive recording tool. It becomes part of a workflow where spoken input can create, clean up, and revise text without requiring constant manual correction.

Custom Vocabulary Could Reduce Manual Fixes

One of the most important additions is customized vocabulary. Users can give the model terms that need to be spelled or handled in a specific way, including unique spelling requirements and specialized jargon.

This is a direct answer to a common transcription problem: general speech models often struggle with words that are obvious to a specific team, field, or project but uncommon in everyday speech. When those words are misheard, a transcript may be technically complete but still frustrating to use.

With customized vocabulary, 3.5 Transcribe can adapt to those terms automatically. The benefit is not just accuracy in a narrow sense. It also reduces the need to search through a transcript afterward and repair the same repeated mistakes.

The model can also attribute speech for up to three speakers in pre-recorded audio. That adds another layer of structure, especially for recordings where knowing who said what is as important as the words themselves.

Word-level timestamps are included as well. That gives users a way to connect specific text back to the original audio, which can help when checking a transcript, reviewing a moment, or moving from written notes back to spoken context.

Gemini 3.5 Live Focuses On Voice Interaction

Gemini 3.5 Transcribe is launching alongside Gemini 3.5 Live and Gemini 3.5 Live Experimental. These models build on the speech recognition technology behind Gemini’s voice chat mode.

Gemini 3.5 Live is designed to improve several parts of real-time voice interaction. Google says it is better at handling mid-sentence interruptions, language recognition, and live visual processing.

Those details point to a broader goal for Gemini Audio: making voice-controlled AI feel less brittle. Real conversations are messy. People pause, restart, talk over their own thoughts, switch context, and speak while other noise is present. A voice system that fails under those conditions is less useful than one that can keep up.

Gemini 3.5 Live Experimental goes further in a different direction. It can narrate its progress step by step in real time while working through reasoning on more complex tasks.

That makes the experimental version less about transcription alone and more about making an AI assistant’s process visible while it is working. The source does not specify the full range of tasks, but it frames the feature around more complex reasoning and real-time progress narration.

Where The Update Is Available

The Gemini Audio updates are rolling out starting today in English for all macOS Gemini app users. They are also coming to the Rambler dictation feature on Android in select countries and languages.

Developers can access the update in public preview through the Gemini API via AI Studio and Antigravity. Google also says Chrome support is coming soon.

The rollout is arriving while Google has still not released the Gemini 3.5 Pro model that it promised to roll out in June. That timing makes the audio update notable for two reasons: it brings new Gemini 3.5 capabilities now, and it arrives before the delayed Pro model mentioned in the source.

For users, the most immediate change is likely to be in everyday dictation and transcript cleanup. Removing filler words, handling interruptions, recognizing more than 85 languages, adapting to jargon, and labeling up to three speakers are all features aimed at turning spoken input into more usable text.

For developers, the public preview means these speech capabilities can start being tested inside products and workflows built with the Gemini API. For everyone else, the practical question is simple: whether Gemini Audio can make voice input less like raw recording and more like structured, editable text from the start.