27 August 2026
Google releases Gemini 3.5 Transcribe speech-to-text model
First reported
Ars Technica and Google DeepMind ran this on , a day before the other 2 sources picked it up.
- The model converts spoken audio into formatted text automatically, removing filler words and correcting speech errors across 85 languages.
- It processes real-time speech 70 percent faster than Google's previous model, Chirp 3, with error rates of 4.0 percent for live streaming and 2.6 percent for recorded audio.
How it was covered
Reported by The Decoder, Ars Technica, Google DeepMind