AMD and Cerebras Partner on Hybrid AI Inference Solution
AMD and Cerebras Systems have announced a partnership to develop a joint AI inference solution that integrates AMD's Helios rack-scale systems with Cerebras' wafer-scale engines. Following the announcement, Cerebras' stock price rose by 11% during intraday trading. This architectural collaboration addresses the distinct phases of LLM inference by splitting the workload, potentially setting a new standard for high-throughput, low-latency AI processing. It also strengthens both companies' positions in the competitive AI hardware market against Nvidia's dominant ecosystem. In the joint workflow, AMD Helios will handle high-throughput prompts and large context windows, while Cerebras' wafer-scale engine will manage memory-bandwidth-intensive token generation. The solution is expected to debut on Cerebras Cloud in the second half of 2026 before a broader rollout.
## BACKGROUND
LLM inference typically consists of two phases: the prefill phase (processing the input prompt) and the decoding phase (generating tokens one by one). AMD's Helios is a rack-scale system combining Instinct GPUs, EPYC CPUs, and Pensando networking to act as a single massive accelerator, while Cerebras' Wafer-Scale Engine is a giant single-chip processor designed to bypass traditional chip packaging limits for ultra-fast memory access.