~/SPEECH RECOG/microsoft-launches-mai-transcribe-2-streaming-with-0-13s-latency-and-2

Microsoft Launches MAI-Transcribe-2-Streaming with 0.13s Latency and 2.5% Word Error Rate

Microsoft has introduced MAI-Transcribe-2-Streaming, a real-time speech-to-text AI model supporting 60 languages with continuous language detection. In Artificial Analysis evaluations, the model ranked first among 28 streaming speech models with a 2.5% word error rate and a latency of 0.13 seconds. Ultra-low latency streaming speech recognition allows interactive voice assistants and live captioning tools to process user intent before a sentence is completed. Achieving top performance across latency and accuracy benchmarks helps establish a new industry standard for real-time AI audio applications. The model outputs initial partial text within around 100 milliseconds of audio input and refines the transcript contextually until finalization. It costs $0.54 per hour (about $9 per 1,000 minutes) and is available via Microsoft Foundry, MAI Playground, and OpenRouter.

## BACKGROUND

Word Error Rate (WER) is the standard metric used to assess speech recognition systems, where a lower percentage indicates greater accuracy. Unlike traditional batch transcription that processes complete audio files, streaming speech-to-text outputs text continuously as speech occurs, making low latency vital for live dialogue.

## REFERENCES

## KEYWORDS

#Speech Recognition#Artificial Intelligence#Microsoft#Real-Time Systems#Audio AI

$ subscribe --daily

Microsoft Launches MAI-Transcribe-2-Streaming with 0.13s Latency and 2.5% Word Error Rate | Daily News