openTPU: Open-Source AI Inference Hardware Engine Designed with LLM Assistance
The openTPU project has introduced an open-source AI hardware inference engine designed with LLM assistance to run modern open-weight models such as Qwen and Gemma. Through human-guided LLM optimization loops, the design evolved to achieve over 80 tokens per second on smaller models. This project highlights the expanding capability of large language models in electronic design automation (EDA) and hardware accelerator development. It demonstrates how AI-assisted iterative optimization can dramatically lower the threshold for creating custom domain-specific silicon architectures. Rather than fully autonomous hardware synthesis, openTPU was created by prompting an LLM within a simulation framework to refine design constraints. Building on methods previously used to create RISC-V CPU cores, the hardware iteratively improved its throughput from a few tokens per second up to 80+ tokens per second.
## BACKGROUND
Electronic Design Automation (EDA) comprises software tools used to design, simulate, and verify integrated circuits. Tensor Processing Units (TPUs) are application-specific integrated circuits (ASICs) tailored specifically to accelerate neural network workload computations.