~/NVIDIA BLACK/nvidia-blackwell-gb300-sets-moe-pre-training-record-with-deepseek-v3

Nvidia Blackwell GB300 Sets MoE Pre-Training Record with DeepSeek-v3

Nvidia announced that its Blackwell GB300 NVL72 platform achieved a record-breaking 1,648 TFLOPs per GPU during the pre-training of the DeepSeek-v3 671B model. This represents a 50% performance improvement over tests conducted in November 2025 and a nearly threefold increase compared to the previous GB200 generation. This milestone demonstrates significant advancements in AI hardware efficiency, allowing organizations to train massive Mixture-of-Experts (MoE) models using fewer hardware resources and lower computational costs. The record was achieved using 256 GPUs to train the 671-billion-parameter model, driven by 38 major system optimizations identified through 1.4 million hours of GPU testing.

## BACKGROUND

Mixture-of-Experts (MoE) is a machine learning architecture that splits a large model into smaller sub-networks, activating only a fraction of them for each token to reduce computational costs. The Nvidia GB300 NVL72 is a liquid-cooled rack-scale system that integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs, designed to act as a single high-performance supercomputer node.

## REFERENCES

## KEYWORDS

#Nvidia Blackwell#DeepSeek-v3#Mixture of Experts (MoE)#AI Hardware#Distributed Training

$ subscribe --daily

Nvidia Blackwell GB300 Sets MoE Pre-Training Record with DeepSeek-v3 | Daily News