Google Launches Gemini 3.5 Transcribe, Its Most Accurate Speech Model
The system detects 85+ languages automatically, tags up to three speakers with timestamps, and streams audio 70% faster than Chirp 3.
- Streaming transcription averages 4.0% word error rate, non-streaming reaches 2.6%, per Artificial Analysis.
- Two separate APIs ship: gemini-3.5-transcribe-live for real time, gemini-3.5-transcribe for recordings.
- Early partners include Agora, LiveKit, LangChain, Vercel and Pipecat testing the models.
Why it matters: A speech model that never forgets what it hears turns conversation into a permanent record held by Google.