27 August 2026

Google releases Gemini 3.5 Transcribe speech-to-text model

First reported

Ars Technica and Google DeepMind ran this on , a day before the other 2 sources picked it up.

  • The model converts spoken audio into formatted text automatically, removing filler words and correcting speech errors across 85 languages.
  • It processes real-time speech 70 percent faster than Google's previous model, Chirp 3, with error rates of 4.0 percent for live streaming and 2.6 percent for recorded audio.

How it was covered

Reported by The Decoder, Ars Technica, Google DeepMind