Moore Threads Achieves Day-0 Adaptation of MiniMax H3 Multimodal Model
Moore Threads has achieved day-0 adaptation and deployment of MiniMax's newly open-sourced H3 multimodal generative model on its MTT S5000 GPU. The adaptation was completed in just three hours using the SGLang-MUSA inference framework and the muDNN operator library. This rapid deployment demonstrates the growing maturity and agility of China's domestic AI hardware and software ecosystems in supporting cutting-edge multimodal models. It also lowers the barrier for enterprises to deploy cost-effective, commercial-grade video and audio generation locally. The MiniMax H3 model supports text, image, audio, and video inputs to generate native audio-video content up to 2K resolution and 15 seconds long. Moore Threads utilized its MUSA software stack, mapping the model to its high-performance MATE and muDNN libraries to ensure stable execution.
## BACKGROUND
MUSA is Moore Threads' proprietary GPU software stack designed as an alternative to NVIDIA's CUDA environment, enabling developers to run and port AI workloads on domestic Chinese GPUs. SGLang is an open-source, high-performance serving framework designed for structured generation and fast inference of large language and multimodal models.