~/AI ML/alibaba-releases-qwen-audio-3-1-suite-and-slashes-api-prices-up

Alibaba Releases Qwen-Audio-3.1 Suite and Slashes API Prices up to 95%

Alibaba's Qwen team has released the Qwen-Audio-3.1 model suite, featuring upgraded speech recognition, text-to-speech, and real-time interaction capabilities alongside next-generation ASR-Next and TTS-Next models. Concurrently, API prices were dramatically slashed across all audio products, with ASR dropping by 95%, Realtime by 85%, and TTS by 70%. This release significantly lowers cost barriers for developers building speech-driven applications, from streaming transcriptions to end-to-end podcast and game audio creation. By delivering multi-speaker voice consistency and combined ambient sound generation at a fraction of previous prices, Alibaba is accelerating the deployment of real-time audio AI across industries. Qwen-Audio-3.1-ASR supports 30 languages and 16 Chinese dialects with a streaming first-token latency of around 160 milliseconds and native text-polishing features. Additionally, Qwen-Audio-3.1-TTS-Next can generate dialogue, sound effects, and background music in a single unified pass while maintaining character voice consistency across long-form interactions.

## BACKGROUND

Audio Large Language Models (Audio LLMs) expand traditional text models by natively understanding and generating acoustic wave signals, bypassing older cascaded pipelines that separate speech recognition, text reasoning, and voice synthesis. Alibaba's foundational Qwen-Audio architecture previously unified diverse speech, music, and sound understanding tasks into a single transformer backbone.

## REFERENCES

## KEYWORDS

#AI/ML#Speech Recognition#Qwen#Audio LLMs#Generative AI

$ subscribe --daily

Alibaba Releases Qwen-Audio-3.1 Suite and Slashes API Prices up to 95% | Daily News