Microsoft Releases Two New MAI Models for Image Generation and Speech
Microsoft’s AI division has released MAI-Image-2.5-Pro for image generation and MAI-Voice-2-Flash for speech synthesis, both available in Azure AI Foundry and the MAI Playground. Image-2.5-Pro is described as Microsoft’s highest-quality image model with strong text rendering inside images ($5/1M input text tokens, $8/1M input image tokens, $106/1M output tokens), while Voice-2-Flash is twice as fast and 32% cheaper than its predecessor at $15/1M characters. Internally, Microsoft is replacing third-party models with its own: Bing Image Creator now runs entirely on MAI-Image-2.5, PowerPoint image editing costs dropped up to 84%, OneDrive saw a 26% increase in saved results, and Voice-2-Flash cut compute costs by 89% in Dynamics 365 Contact Center.