Alibaba Launches Qwen3.8-Omni-Flash Native Omnimodal Model with 1M Context Window
Alibaba officially launched Qwen3.8-Omni-Flash, a native omnimodal AI model capable of processing text, image, audio, and video inputs within a 1-million-token context window. Across 30 evaluation benchmarks, the new model achieves an average score improvement of over 26% compared to Qwen3.5-Omni-Plus, while reducing audio and video API costs by over 90%. The model brings high-performance multimodal understanding—including real-time spatial audio perception and complex video-audio post-production—to enterprise applications at drastically reduced inference costs. By matching or exceeding models like Gemini 3.8 Flash in key audio and agentic tasks while cutting token consumption by 45.7%, it sets a new efficiency benchmark for real-time omnimodal AI. Qwen3.8-Omni-Flash introduces a Realtime version supporting spatial audio localization to determine a sound source's direction and distance alongside video streams. Additionally, Alibaba open-sourced the Qwen-Live Harness framework for continuous multimodal streaming and expanded Qwen-MM-Plugins, featuring utility tools such as Video2Note for generating formatted PDF notes from multi-hour videos.
## BACKGROUND
Omnimodal AI models process multiple media types—such as text, vision, and speech—within a single unified neural network rather than using separate cascading models. Evaluating these advanced systems requires specialized speech metrics like cpWER (Concatenated Permutation Word Error Rate) for multi-speaker transcription, as well as complex agent benchmarks like AgenticVBench and WildClawBench-MM that measure long-horizon task execution.