Microsoft Ships Its First Streaming Transcription Model and Multilingual Voice Models
Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming transcription model, alongside two speech synthesizers: MAI-Voice-2.1 and the lighter MAI-Voice-2.1-Flash. The transcriber covers 60 languages with automatic continuous language detection, returns its first partial hypotheses in just over 100ms of incoming audio, and ranks no. 1 for accuracy on both final and partial transcripts in Artificial Analysis evaluations, at an introductory price of $0.54 per hour of audio through the end of the year. MAI-Voice-2.1 supports 23 languages and 26 locales, with a single voice speaking each language in a native accent instead of carrying one accent across all of them, at $22 per 1M characters, while Flash generates up to 45s of audio at 150ms end-to-end latency for $15 per 1M characters. Both voice models clone a speaker from a few seconds of reference audio and ship with consent guardrails, and all three models are available in Microsoft Foundry, the MAI Playground, Azure Voice Live and Vercel, with OpenRouter carrying the voice models and a live voice agent demo called Chatter in the MAI Playground.
Related: Microsoft Releases Two New MAI Models for Image Generation and Speech, Microsoft AI Launches Seven New MAI Models
Our first streaming transcription model debuts at no. 1 on Artificial Analysis