A 5.2KB Pure x86-64 Assembly Inference Engine for Gemma-2B
Developer tom_tsai28 released PULSAR-ASM, an open-source bare-metal Gemma-2B inference engine written entirely in flat x86-64 assembly using FASM. The 5.2KB binary achieves around 4.6 tokens per second in FP16 precision on an older quad-core Intel i5 CPU without any C/C++ or PyTorch runtime dependencies. This project demonstrates the minimal memory footprint required to run modern autoregressive Transformer models directly on CPU hardware without heavy software runtimes. It offers useful architectural insights and a baseline reference for deploying micro-LLMs on resource-constrained microcontrollers and DSPs. The engine utilizes AVX2 vector instructions and F16C extensions for half-precision floating-point conversions, incorporating a custom 4-thread SMP GEMM implementation that sustains ~18.5 GB/s memory bandwidth on DDR4-2400 RAM. A lightweight Python harness interacts with the binary strictly via `ctypes` for OS memory allocation and thread creation.
## BACKGROUND
Modern Large Language Model (LLM) inference usually relies on massive software stacks like PyTorch, CUDA, or C++ runtimes to manage matrix multiplications (GEMM) and tensor operations. In contrast, x86-64 assembly language enables developers to write direct machine instructions for the CPU, leveraging instruction set extensions like AVX2 for parallel vector processing and F16C for hardware-accelerated 16-bit floating-point conversions.