Intelligent transcription with Gemini 3.5 Transcribe
Google is putting Gemini 3.5 Transcribe into public preview for developers, with separate live and recorded-audio APIs.
The model is pitched as Google’s most precise speech-to-text system yet, handling filler-word cleanup, self-corrections, formatting, custom vocabulary, and more than 85 languages. Google says it reaches 4.0% WER for streaming and 2.6% for non-streaming in Artificial Analysis measurements, with a 70% improvement in time to final transcription over Chirp 3. It is also already showing up in Google products including Rambler on Android, the Gemini app on macOS, Gboard, Antigravity, and later Chrome. Google DeepMind's note
The model is pitched as Google’s most precise speech-to-text system yet, handling filler-word cleanup, self-corrections, formatting, custom vocabulary, and more than 85 languages. Google says it reaches 4.0% WER for streaming and 2.6% for non-streaming in Artificial Analysis measurements, with a 70% improvement in time to final transcription over Chirp 3. It is also already showing up in Google products including Rambler on Android, the Gemini app on macOS, Gboard, Antigravity, and later Chrome. Google DeepMind's note
score 6