DeepSeek Open-Sources AI Infrastructure Software Stack for Huawei Ascend Hardware
DeepSeek has officially open-sourced its core AI infrastructure components ported for Huawei's Ascend platform, including the TileLang domain-specific language compiler and high-performance libraries such as DeepGEMM, DeepEP, FlashMLA, TileKernels, and DeepSelect. These components directly mirror their NVIDIA CUDA counterparts and achieve near-hardware-limit performance on Ascend hardware. By bringing high-performance LLM training and inference capabilities to Huawei Ascend hardware, DeepSeek reduces the AI industry's reliance on NVIDIA GPUs and provides a viable software stack for alternative hardware. This bridges the optimization software gap for domestic Chinese AI accelerators, facilitating efficient training of massive models like DeepSeek V4 on non-NVIDIA infrastructure. TileLang encapsulates Huawei's low-level Ascend C instructions into a high-level domain-specific language, enabling simpler code logic and faster kernel development without sacrificing execution speed. The accompanying open-source libraries optimize key AI workloads: DeepGEMM handles matrix operations, DeepEP accelerates cross-device communication, FlashMLA improves sparse attention efficiency in long contexts, TileKernels handles vector computations, and DeepSelect speeds up data filtering.
## BACKGROUND
Developing custom operators for specialized AI chips like Huawei Ascend typically relies on low-level languages like Ascend C (built on C/C++), which require significant engineering effort to reach hardware limits compared to NVIDIA's mature CUDA ecosystem. TileLang is a domain-specific language (DSL) designed to streamline GPU and NPU kernel programming using a composable tiled programming model. Furthermore, FlashMLA is DeepSeek's optimized attention library tailored for Multi-head Latent Attention (MLA), an architecture crucial for efficient context handling in DeepSeek models.