~/NVIDIA/nvidia-reportedly-restarts-and-redesigns-rubin-cpx-gpu-for-ai-prefill

Nvidia Reportedly Restarts and Redesigns Rubin CPX GPU for AI Prefill

Nvidia has reportedly restarted its Rubin CPX GPU project with a major redesign aimed at accelerating the prefill phase of AI inference, targeting mass production in Q1 2027. The updated architecture swaps its memory design from 128GB GDDR7 to 168GB HBM4 while increasing compute performance close to standard Rubin GPUs at a max power consumption of 2,300W. As long-context LLM workloads scale, prompt prefill and key-value (KV) cache building account for over 50% of inference resources. Deploying prefill-specialized hardware like Rubin CPX enables data centers to scale inference infrastructure more flexibly and cost-effectively by disaggregating prefill workloads from decode tasks. Eight Rubin CPX chips form a compute tray delivering 1.34TB of HBM memory, connected via Spectrum-6 Ethernet with RDMA links to standard Vera Rubin NVL72 racks at a recommended 1:1 ratio. Switching from GDDR7 to HBM4 underscores the immense memory bandwidth required to handle long-context prefill workloads efficiently.

## BACKGROUND

Large language model (LLM) inference is split into two phases: the prefill phase, which processes the input prompt in parallel to compute the key-value (KV) cache, and the decode phase, which generates output tokens one by one autoregressively. Because prefill is compute- and bandwidth-intensive while decode is memory-bound, separating them onto dedicated hardware architectures improves hardware utilization and latency.

## REFERENCES

## KEYWORDS

#Nvidia#AI Hardware#GPU Architecture#LLM Inference#HBM4

$ subscribe --daily

Nvidia Reportedly Restarts and Redesigns Rubin CPX GPU for AI Prefill | Daily News