Huawei Unveils NPO 4096-Card Supernode, Accelerates Ascend AI Roadmap, and Highlights openEuler Lead
Huawei announced an accelerated roadmap for its Ascend 960/970/980 AI chips, introduced the industry's first NPO-based 4096-card supernode offering up to 8 EFLOPS FP8 compute capacity, and revealed that openEuler has surpassed 20 million deployments to secure the top server OS market share in China. The accelerated chip timeline and NPO supernode architecture highlight Huawei's commitment to building self-reliant, ultra-large-scale AI infrastructure for massive AI model training. Furthermore, PyTorch natively supporting Ascend marks a major milestone in global developer adoption for Huawei's AI compute platform. The Ascend 960 supernode aggregates 4,096 GPUs with 1 PB of HBM capacity, utilizing 5,500 Hi-ONE modules to replace 48,000 800G optical transceivers—cutting power consumption by over 550 kW and doubling system uptime. Additionally, external developers now make up 61% of the CANN software ecosystem, outnumbering Huawei's internal developers for the first time.
## BACKGROUND
Near-Packaged Optics (NPO) is an advanced optical interconnect architecture where optical engines are mounted on the PCB directly adjacent to high-density compute or switch ASICs, significantly cutting latency and power usage compared to traditional pluggable transceivers. Huawei's Lingqu UnifiedBus is a high-speed interconnect protocol engineered to bridge thousands of discrete AI accelerators so they function logically as a single unified supercomputer.