Microsoft AI is claiming a multilingual speech record for its newest transcription model. MAI-Transcribe-2 averaged a 5.2% word error rate across 60 languages on the FLEURS benchmark, a result the company says tops every rival system tested.
Microsoft prices the model at $0.10 for every hour of audio it processes and calls it the most capable transcription system it has built. Clinical note-taking and legal documentation are the early targets, along with accessibility tools and closed captioning.
In Microsoft’s telling, MAI-Transcribe-2 outperforms Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2 across a broad mix of real-world audio. The lab cites Artificial Analysis rankings that place the model on the accuracy-latency Pareto frontier and second on the word-error leaderboard.
Microsoft also claims up to 10 times the processing speed of leading rivals on long-form audio, with quality that holds in noisy conditions.
The launch arrives in a crowded week for speech models and adds weight to Microsoft’s in-house MAI family, which the company has been offering as an alternative to other labs’ models.