1 link tagged with all of: gemini-api + real-time-transcription + speech-to-text + developer-tools + voice-ai
Links
Google launched two new speech models for developers: Gemini 3.8 Live handles real-time voice conversations with reasoning and background tool execution, while Gemini 3.5 Transcribe converts speech to text across 85+ languages with a 4.0% error rate.
- Gemini 3.8 Live can execute API calls in the background while streaming audio responses, handle visual context, and support 97+ languages with accent consistency
- Gemini 3.5 Transcribe achieves 4.0% word error rate (streaming) and 2.6% (non-streaming), with automatic code-switching and custom vocabulary biasing for domain-specific terms
- Extended Thinking variant adds multi-step reasoning capabilities, ranking #1 on Artificial Analysis' Speech-to-Speech leaderboard
- Google's full audio suite includes speech translation (70+ languages), text-to-speech, and music generation all available in the Gemini API
voice-ai
speech-to-text
gemini-api
real-time-transcription
developer-tools