~/AI HARDWARE/alibaba-cloud-s-zhenwu-m890-supernode-runs-2-4-trillion-parameter-qwen3

Alibaba Cloud's Zhenwu M890 Supernode Runs 2.4-Trillion-Parameter Qwen3.8 Model

Alibaba Cloud announced that its Lingjun Zhenwu M890 supernode has successfully adapted to the 2.4-trillion-parameter Qwen3.8 model, making it China's first supernode to run a model exceeding 2 trillion parameters. The model is now available on Alibaba Cloud's Model Studio (Bailian) platform for inference services. This achievement marks a significant milestone in large-scale AI system engineering, demonstrating China's capability to run massive Mixture of Experts (MoE) models on domestic hardware. By optimizing the full link from chip to cloud, it lowers inference costs and improves performance by up to 1.5 times in Agentic scenarios. The Zhenwu M890 is a next-generation AI chip developed by T-Head that natively supports data precisions from FP32 to FP4. Using the ICN Switch 1.0 interconnect chip, a single instance connects 64 Zhenwu M890 chips with an 800GB/s bandwidth, providing 9TB of video memory to support expert parallelism.

## BACKGROUND

Mixture of Experts (MoE) is an AI architecture that activates only a subset of the network (experts) for each input, allowing models to scale to trillions of parameters. Expert parallelism is a technique used to distribute these different experts across multiple GPUs to overcome memory limitations. Standard AI clusters often struggle with the high communication bandwidth required to keep these distributed parameters synchronized.

## REFERENCES

## KEYWORDS

#AI Hardware#Large Language Models#System Architecture#Alibaba Cloud

$ subscribe --daily

Alibaba Cloud's Zhenwu M890 Supernode Runs 2.4-Trillion-Parameter Qwen3.8 Model | Daily News