⏳ Curating articles…
Artificial Intelligence 2 min read 2h ago

Google Gemini Scrubs Ums from Voice Text

  • Gemini 3.5 Transcribe, a new model, automatically removes filler words, supports 85-plus languages, and can attribute speech to up to three speakers in pre-recorded audio.
  • Gemini 3.5 Live improves on mid-sentence interruption handling and live visual processing, while the Experimental variant narrates its reasoning steps in real time.
  • The models are rolling out today for macOS Gemini app users and Android Rambler in select markets, with Chrome support still to come and the previously promised Gemini 3.5 Pro yet
Google Gemini Scrubs Ums from Voice Text

Google has updated its Gemini Audio suite with three new speech models — Gemini 3.5 Live, Gemini 3.5 Live Experimental, and Gemini 3.5 Transcribe — introducing capabilities that automatically remove filler words, detect specialised jargon, and support more than 85 languages.

A new transcription model

Gemini 3.5 Transcribe is an entirely new addition to the Gemini family. Google describes it as representing "a major advancement" over its previous transcription model, Chirp 3, particularly in multilingual performance and wording error rates. The model can automatically format text, strip out filler words such as "um" and "uh", and allow users to supply a customised vocabulary so that specialist terminology and unconventional spellings are handled without manual correction. It can also attribute speech to up to three distinct speakers in pre-recorded audio and generate word-level timestamps.

Live models and real-time reasoning

Gemini 3.5 Live builds on the speech recognition technology that underpins Gemini's existing voice chat mode, with improvements to mid-sentence interruption handling, language recognition, and live visual processing. Gemini 3.5 Live Experimental extends this further, narrating its reasoning progress step by step in real time when tackling more complex tasks.

Advertisement
Ad Unit · 728×90 / Responsive

Availability and what remains unresolved

The Gemini Audio updates are rolling out today in English for all macOS Gemini app users and via the Rambler dictation feature on Android in select countries and languages. Developers can access the models in public preview through the Gemini API via AI Studio and Antigravity, with Chrome support described as coming soon. The launch arrives as Google's separately promised Gemini 3.5 Pro model, which the company indicated would be released in June, has yet to appear — a gap the company has not publicly addressed. The question of how effectively the new transcription model handles highly technical or low-resource languages beyond those demonstrated also remains open.

Advertisement
Ad Unit · 300×250 / Responsive

More in Artificial Intelligence

Read in another language

← Home