Nvidia Blackwell GB300 Sets MoE Pre-Training Record with DeepSeek-v3
Nvidia announced that its Blackwell GB300 NVL72 platform achieved a record-breaking 1,648 TFLOPs per GPU during the pre-training of the DeepSeek-v3 671B model. This represents a 50% performance improvement over tests conducted in November 2025 and a nearly threefold increase compared to the previous GB200 generation. This milestone demonstrates significant advancements in AI hardware efficiency, allowing organizations to train massive Mixture-of-Experts (MoE) models using fewer hardware resources and lower computational costs. The record was achieved using 256 GPUs to train the 671-billion-parameter model, driven by 38 major system optimizations identified through 1.4 million hours of GPU testing.
## BACKGROUND
Mixture-of-Experts (MoE) is a machine learning architecture that splits a large model into smaller sub-networks, activating only a fraction of them for each token to reduce computational costs. The Nvidia GB300 NVL72 is a liquid-cooled rack-scale system that integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs, designed to act as a single high-performance supercomputer node.