More on the topic…
Microsoft just released MAI-Transcribe-2, a speech recognition model priced at 10 cents per hour—a 72% price cut from the first version released five months ago at $0.36. For a large bank processing 100,000 hours of call-center audio yearly, that drops costs from $36,000 to $10,000, making transcription essentially free as a business expense. The model handles 60 languages (up from 43 in June's update and 25 in April's original), and Microsoft bundled in features competitors typically charge extra for: speaker diarization, word-level timestamps, keyword biasing, automatic language detection, and code-switching for conversations that mix languages mid-sentence like Hinglish or Spanglish. There's also a verbatim mode for legal compliance and a clean mode for readable captions—all included at the base price.
Microsoft's performance claims rest on three different benchmarks, each telling a different story. On FLEURS (Google's multilingual benchmark), it scores 5.2% word error rate across 60 languages, ranking first—though that average actually rose from 3.7% on the previous version because averaging across more languages includes low-resource ones where accuracy naturally drops. On Artificial Analysis's leaderboard, which tests models through their actual public APIs, it ranks second and defines the accuracy-latency Pareto frontier, meaning no competitor beats it on speed without sacrificing accuracy or vice versa. The speed claims are real: 10 times faster than OpenAI's GPT-Transcribe, seven times faster than ElevenLabs' Scribe v2, five times faster than Google's Gemini 3.5. That efficiency matters because batch transcription cost depends on GPU hours, not waiting time—a model running 300 times real-time uses a fraction of the compute of one at 30 times.
Three releases in five months reveals Microsoft's actual strategy. MAI-Transcribe-1 launched April 2, MAI-Transcribe-1.5 on June 2, and MAI-Transcribe-2 today—each expanding language coverage by roughly 40% while adding premium features at commodity prices. This cadence signals a team with a stable architecture now scaling data and improving predictably, exactly the phase where speech models improve fastest. Mustafa Suleyman, who leads Microsoft AI after joining from Inflection in March 2024, told The Verge the original team was just 10 people "liberated from bureaucracy," running at half the GPU cost of competitors. That speed and independence matter because Microsoft is systematically replacing OpenAI technology in its own products. Despite investing $13 billion in OpenAI, Salesforce CEO Marc Benioff predicted in January that Microsoft will eventually use its own frontier models instead, which is precisely what's happening with transcription.
Questions about this article
No questions yet.