Alibaba Launches Qwen-Audio-3.0 Speech Models on Qianwen AI Platform
Alibaba Cloud has officially launched its Qwen-Audio-3.0 series of speech models on the Qianwen AI and Bailian platforms, accessible via API and Token Plan. The series includes three models: Qwen-Audio-3.0-ASR for speech recognition, Qwen-Audio-3.0-TTS for speech synthesis, and Qwen-Audio-3.0-Realtime for real-time voice interaction. The Qwen-Audio-3.0 series achieved the top global ranking in ASR, RealTime, and TTS categories on the Artificial Analysis leaderboard, marking a major advancement in speech AI. This release enhances Alibaba's competitiveness in end-to-end speech-to-speech dialogue systems, which are crucial for natural, low-latency human-AI interaction. Qwen-Audio-3.0-Realtime supports end-to-end speech understanding, allowing users to interrupt and ask follow-up questions while also supporting tool calling. Meanwhile, the TTS model allows control over emotion, tone, and rhythm, and the ASR model is optimized for complex contexts and professional domains.
## BACKGROUND
Traditional voice assistants rely on a cascaded pipeline: converting speech to text (ASR), processing it with a language model, and converting the text response back to speech (TTS). In contrast, modern end-to-end speech-to-speech systems directly map spoken input to spoken output, significantly reducing latency and preserving nonverbal cues like emotion. Artificial Analysis is an independent benchmarking platform that evaluates AI models on performance, quality, and cost.