~/AI HARDWARE/astera-labs-announces-leo-x-and-leo-2-smart-memory-controllers-for

Astera Labs Announces Leo X and Leo 2 Smart Memory Controllers for LLM Offloading

Astera Labs introduced its Leo X, Leo 2 E, and Leo 2 P series smart memory controllers supporting CXL 3.2 and PCIe Gen6 specifications. These controllers enable AI workload infrastructure to offload KV Cache and agent context to external memory tiers, reducing time-to-first-token by up to 62% and increasing token throughput by up to 22%. Limited GPU VRAM capacity is a primary bottleneck for serving Large Language Models (LLMs) with long context windows and high concurrency. By providing dedicated, low-latency external memory access via CXL 3.2 and PCIe Gen6, data centers can scale AI inference memory capacity cost-effectively without relying solely on expensive onboard GPU memory. The Leo 2 E and Leo 2 P chips support four downstream DDR4/DDR5 memory controllers, with the Leo 2 P featuring a 2x8 dual-port configuration for memory pooling and sharing across hosts. The fabric-based Leo X series works alongside Astera's Scorpio network switch chips to scale out GPU memory pools over PCIe and custom accelerator protocols.

## BACKGROUND

During LLM inference, key-value (KV) caching stores intermediate attention tensors to prevent recomputing past tokens, but long context lengths quickly consume available GPU VRAM. Compute Express Link (CXL) is an open industry interconnect standard that enables low-latency, cache-coherent memory expansion and sharing between CPUs, GPUs, and external memory controllers over PCIe links.

## REFERENCES

## KEYWORDS

#AI Hardware#CXL#KV Cache#Astera Labs#LLM Infrastructure

$ subscribe --daily

Astera Labs Announces Leo X and Leo 2 Smart Memory Controllers for LLM Offloading | Daily News