~/HARDWARE/ddr4-pcie4-vs-ddr5-pcie5-benchmarks-for-llm-pre-training

DDR4/PCIe4 vs DDR5/PCIe5 Benchmarks for LLM Pre-Training

A user benchmarked LLM pre-training workloads across DDR4/PCIe4 and DDR5/PCIe5 platforms, showing a 15–20% throughput boost on the newer hardware. However, due to current high RAM prices, allocating budget toward an additional GPU on a DDR4/PCIe4 system yields up to 50% more throughput for roughly the same cost. This empirical comparison helps AI developers optimize their hardware budgets when building workstations for deep learning. It demonstrates that adding extra GPU compute power generally offers far better performance per dollar than upgrading host RAM and PCIe interconnect speeds for training tasks. The tests compared an AMD EPYC 7352 node (PCIe4 at 26.3 GB/s, 192 GB DDR4) against an AMD Threadripper 9975WX node (PCIe5 at 54.3 GB/s, 256 GB DDR5 @ 6400 MT/s) using Vast.ai rentals. Caveats of choosing older DDR4 platforms include limited motherboard replacement availability and potential BIOS/POST incompatibilities with future GPU architectures.

## BACKGROUND

System memory speed (measured in MT/s) and PCIe bus bandwidth control how fast data transfers between the CPU host and GPUs. While LLM inference and multi-GPU tensor parallelism often rely heavily on host-to-device bandwidth, pre-training workloads are predominantly compute-bound by GPU processing power.

## REFERENCES

## KEYWORDS

#Hardware#LLM#Benchmarking#PCIe#AI Workstations

$ subscribe --daily

DDR4/PCIe4 vs DDR5/PCIe5 Benchmarks for LLM Pre-Training | Daily News