1 link tagged with all of: developer-tools + voice-ai + gemini-api + speech-to-text
Click any tag below to further narrow down your results
Links
Google launched two new speech models for developers: Gemini 3.8 Live handles real-time voice conversations with reasoning and background tool execution, while Gemini 3.5 Transcribe converts speech to text across 85+ languages with a 4.0% error rate.
- Gemini 3.8 Live can execute API calls in the background while streaming audio responses, handle visual context, and support 97+ languages with accent consistency
- Gemini 3.5 Transcribe achieves 4.0% word error rate (streaming) and 2.6% (non-streaming), with automatic code-switching and custom vocabulary biasing for domain-specific terms
- Extended Thinking variant adds multi-step reasoning capabilities, ranking #1 on Artificial Analysis' Speech-to-Speech leaderboard
- Google's full audio suite includes speech translation (70+ languages), text-to-speech, and music generation all available in the Gemini API