~/GPU/moore-threads-adapts-mtt-s5000-gpu-to-moonshot-ai-s-2-8t

Moore Threads Adapts MTT S5000 GPU to Moonshot AI's 2.8T Parameter Kimi K3

Moore Threads has announced immediate "Day-0" support for Moonshot AI's newly open-sourced 2.8 trillion parameter Kimi K3 model using its MTT S5000 GPU and MUSA software stack. The adaptation covers model structure parsing, weight loading, core operators, and distributed execution. This achievement showcases the rapid adaptability of domestic Chinese AI hardware and software ecosystems to handle massive, complex hybrid model architectures. It reduces reliance on foreign hardware by proving that domestic GPUs can support cutting-edge, multi-trillion-parameter Mixture-of-Experts (MoE) models. The MTT S5000 GPU features 1000 TFLOPS of dense AI compute and 80GB of VRAM. To support Kimi K3's unique architecture, Moore Threads adapted Triton MUSA for Kimi Delta Attention (KDA) and integrated DeepEP to handle token dispatch and communication across its 896-expert MoE structure.

## BACKGROUND

Kimi Delta Attention (KDA) is a hardware-optimized linear attention mechanism that improves memory efficiency and reduces KV cache usage. MUSA (Meta-computing Unified System Architecture) is Moore Threads' proprietary software stack designed as a CUDA alternative, while DeepEP is a specialized communication library optimized for Mixture-of-Experts (MoE) models.

## REFERENCES

## KEYWORDS

#GPU#LLM#AI Hardware#Kimi K3#Mixture of Experts

$ subscribe --daily

Moore Threads Adapts MTT S5000 GPU to Moonshot AI's 2.8T Parameter Kimi K3 | Daily News