27 August 2026
Google releases Gemini 3.5 Transcribe speech-to-text model
First reported
Ars Technica and Google DeepMind ran this on , a day before the other 2 sources picked it up.
- The model converts spoken audio into formatted text automatically, removing filler words and correcting speech errors across 85 languages.
- It processes real-time speech 70 percent faster than Google's previous model, Chirp 3, with error rates of 4.0 percent for live streaming and 2.6 percent for recorded audio.
- The model is now live in Gboard for Android, the Gemini app on macOS, and Google AI Studio, with Chrome browser support coming soon.
- Developers can build voice applications using two separate APIs: one for real-time streaming and one for recorded audio with speaker identification.
How it was covered
TLDR AITLDR editorial team
Gemini 3.5 Transcribe converts raw audio into accurate, polished, and formatted text for intelligent voice interactions. The model is available on Gemini API and Google AI Studio with support for real-time streaming and pre-recorded audio.
Reported by The Decoder, Ars Technica, Google DeepMind