Google releases Gemini 3.5 Transcribe in public preview
The speech-to-text model is available in Google AI Studio, the Gemini API, and Gemini Enterprise Agent Platform for streaming and pre-recorded audio across 85+ languages.
- Reaches a 4.0% Word Error Rate for streaming and 2.6% for non-streaming audio, improving time to final transcription by 70% compared to Chirp 3
- Provides gemini-3.5-transcribe-live for sub-second streaming and gemini-3.5-transcribe for pre-recorded audio with timestamps
- Automatically filters filler words, formats output, and handles spoken self-corrections across regional accents and noisy environments
- Identifies and attributes speech across up to three distinct speakers, with experimental support for additional speakers
- Rolls out across Google products including Gboard's Rambler on Android, the macOS Gemini app, and Google Antigravity
Developers building voice agents and transcription pipelines gain higher-accuracy speech recognition that natively cleans disfluencies without auxiliary formatting pipelines.

Sources
Read this as text
Back to the AI news