Microsoft Releases Self-Developed MAI-Image-2.5-Pro and MAI-Voice-2-Flash AI Models
Microsoft has launched the public preview of two proprietary AI models, MAI-Image-2.5-Pro for high-quality image generation and MAI-Voice-2-Flash for low-latency voice interaction. These models are trained from scratch using clean, enterprise-grade data without relying on third-party model distillation. The release marks Microsoft's shift toward building independent, cost-effective AI models, significantly reducing GPU operational costs by up to 89% in enterprise applications like PowerPoint and Dynamics 365. It allows Microsoft to offer tailored, enterprise-grade solutions directly integrated into its ecosystem without relying heavily on external partners. MAI-Image-2.5-Pro features improved text-in-image rendering and natural language editing, priced at $106 per million tokens for image output. Meanwhile, MAI-Voice-2-Flash is twice as fast as its predecessor, reducing costs by 32% to $15 per million characters, and is already deployed in Dynamics 365 Contact Center.
## BACKGROUND
Knowledge distillation is a machine learning technique where a smaller, more efficient "student" model is trained to reproduce the behavior of a larger, pre-trained "teacher" model. By avoiding distillation from third-party models, Microsoft ensures full ownership, traceability, and compliance of the training data. This development is supported by Microsoft's newly operational GB200 computing clusters, which power their next-generation AI research.