OpenAI Debuts Custom Jalapeño Inference Chip, Outperforming NVIDIA's GB300 in Efficiency
OpenAI has unveiled Jalapeño, its first custom AI inference chip co-developed with Broadcom, which reportedly delivers 1.5 to 1.9 times the throughput per watt of NVIDIA's GB200 and GB300 systems. Tested using SemiAnalysis's InferenceX benchmark, the chip also reduced end-to-end latency to between 28% and 59% of the NVIDIA systems. This custom silicon milestone allows OpenAI to reduce its heavy reliance on NVIDIA GPUs and lower the massive operational costs associated with running large language models. It signals a broader industry shift where major AI labs design proprietary hardware tailored specifically to their own workloads. Jalapeño is a 700-watt ASIC utilizing a systolic array architecture and HBM, manufactured on TSMC's 3nm process, and is designed strictly for internal inference workloads rather than commercial sale. The chip was developed in a rapid nine-month cycle with the assistance of OpenAI's own AI models, maintaining a sustained power draw of 550 watts or less during testing.
## BACKGROUND
AI inference, the process of running trained models to generate predictions or answers, requires massive computational power and energy. While general-purpose GPUs from NVIDIA currently dominate this space, custom ASICs (Application-Specific Integrated Circuits) can be optimized for specific mathematical operations to achieve superior energy efficiency and lower latency.