~/AI HARDWARE/mlperf-inference-6-1-released-amd-scales-512-mi355x-gpus-as-nvidia

MLPerf Inference 6.1 Released: AMD Scales 512 MI355X GPUs as NVIDIA Vera Rubin Debuts

MLCommons has released the MLPerf Inference v6.1 benchmark results, highlighted by AMD scaling a cluster of 512 Instinct MI355X GPUs to process DeepSeek R1 at up to 2.9 million tokens per second. The benchmark release also featured the first public benchmark submissions for NVIDIA's next-generation Vera Rubin (VR200) architecture. These results demonstrate that AMD can successfully run massive 512-GPU clusters for complex reasoning models like DeepSeek R1, challenging NVIDIA's dominance in large-scale inference deployment. Furthermore, early figures for NVIDIA's Vera Rubin indicate generational performance leaps over current Blackwell-based systems. AMD's 512-GPU MI355X cluster achieved 2,901,950 tokens/s in offline mode and 2,405,310 tokens/s in server mode on DeepSeek R1. In comparison tests, NVIDIA's Vera Rubin VR200 performed 95% faster than a GB300 setup under identical 72-GPU configurations, with a 36-GPU VR200 system even surpassing a 72-GPU GB200 deployment.

## BACKGROUND

MLPerf Inference is an industry-standard benchmark suite maintained by MLCommons that measures hardware throughput and latency across different deployment scenarios such as offline batching and interactive server mode. AMD's Instinct MI355X is built on the CDNA architecture optimized for AI workloads, while NVIDIA's Vera Rubin platform represents its flagship successor architecture to the Blackwell (GB200/GB300) series.

## REFERENCES

## KEYWORDS

#AI Hardware#MLPerf#AMD Instinct#NVIDIA Vera Rubin#DeepSeek

$ subscribe --daily

MLPerf Inference 6.1 Released: AMD Scales 512 MI355X GPUs as NVIDIA Vera Rubin Debuts | Daily News