Google has released Gemini 3.5 Transcribe, ranking #5 on AA-WER at 2.6%, alongside Gemini 3.5 Transcribe Live, achieving 4.0% AA-WER Streaming at 0.40s after speech end
@GoogleAI's Gemini 3.5 Transcribe release comprises two API offerings: Gemini 3.5 Transcribe Live for continuous streaming through the Live API, and Gemini 3.5 Transcribe for pre-recorded audio through the Interactions API. Both support 85+ languages, custom vocabulary and automatic text formatting, while the pre-recorded API adds timestamped multi-speaker identification for up to three speakers, with support beyond three currently experimental.
Key takeaways ➤ Non-streaming transcription: Gemini 3.5 Transcribe achieves 2.6% AA-WER, ranking #5 overall, and processes audio at approximately 84× realtime. ➤ First Final Transcription: Gemini 3.5 Transcribe Live achieves 4.0% WER, with its first final-denoted transcript arriving 0.40s after VAD-detected end of speech. ➤ First Partial Transcription: Gemini 3.5 Transcribe Live achieves 5.8% WER, with its first transcript-bearing event arriving 0.25s after detected end of speech. ➤ Price: Gemini 3.5 Transcribe costs approximately $5 per 1,000 minutes and Transcribe Live $9, assuming 25 audio tokens per second and 175 text tokens per minute at their respective API rates ($2 per 1M audio input tokens and $12 per 1M text output tokens for Transcribe; $3.50 and $21, respectively, for Live).
See more details below ⬇️