Microsoft Launches MAI-Transcribe-2-Streaming with 0.13s Latency and 2.5% Word Error Rate
Microsoft has introduced MAI-Transcribe-2-Streaming, a real-time speech-to-text AI model supporting 60 languages with continuous language detection. In Artificial Analysis evaluations, the model ranked first among 28 streaming speech models with a 2.5% word error rate and a latency of 0.13 seconds. Ultra-low latency streaming speech recognition allows interactive voice assistants and live captioning tools to process user intent before a sentence is completed. Achieving top performance across latency and accuracy benchmarks helps establish a new industry standard for real-time AI audio applications. The model outputs initial partial text within around 100 milliseconds of audio input and refines the transcript contextually until finalization. It costs $0.54 per hour (about $9 per 1,000 minutes) and is available via Microsoft Foundry, MAI Playground, and OpenRouter.
## BACKGROUND
Word Error Rate (WER) is the standard metric used to assess speech recognition systems, where a lower percentage indicates greater accuracy. Unlike traditional batch transcription that processes complete audio files, streaming speech-to-text outputs text continuously as speech occurs, making low latency vital for live dialogue.