More on the topic…
Microsoft's MAI-Transcribe-2 claims the top spot in speech recognition across three metrics: speed, accuracy, and cost. The model hits a 5.2% word-error rate on the FLEURS benchmark across 60 languages and leads the Pareto frontier for accuracy-latency tradeoffs according to Artificial Analysis. It outpaces competitors like OpenAI's GPT-Transcribe (10x faster), ElevenLabs' Scribe v2 (7x faster), and Google's Gemini 3.5 Transcribe (5x faster) while maintaining higher accuracy. The model handles real-world audio that other systems struggle with and supports code-switching between language pairs like Hinglish and Spanglish.
The feature set targets practical applications across legal, medical, and accessibility work. Speaker diarization attributes words to the correct person in multi-speaker recordings. Word-level timestamps enable precise alignment and navigation. Keyword biasing helps catch domain-specific terms and abbreviations. Two transcription modes give developers control: "verbatim" preserves filler words and false starts for compliance, while "clean" removes them for readable captions and notes. The model also automatically detects languages and performs well in noisy environments without needing pre-configuration.
Pricing sits at $0.10 per hour through the end of the year as a limited-time launch offer. The efficiency gains translate directly to cost savings—fewer computing resources needed means lower bills for processing large audio volumes. Access is available now through Microsoft Foundry, MAI Playground, and Open Router for anyone wanting to test it.
Questions about this article
No questions yet.