Custom CUDA Megakernel Accelerates 27B Model to 140 tok/s on Single RTX 3090
A developer released an open-source CUDA megakernel that executes full speculative decoding cycles in a single kernel launch, allowing a 27B parameter model to reach up to 140 tokens/second on an RTX 3090. The custom engine delivers 1.4x to 1.9x faster performance compared to llama.cpp across code generation and prompt processing tasks. This project shows how extreme GPU kernel fusion can unlock massive inference speedups on consumer hardware by eliminating launch overhead and memory bottlenecks. It demonstrates that targeted hardware-and-model optimizations can significantly outperform established general-purpose LLM runtimes like llama.cpp. The engine (open-jet) provides an OpenAI-compatible API server, but it is currently specialized strictly for Unsloth's Q4_K_M quantization on RTX 3090 GPUs. By executing entire draft-and-verify cycles in one launch, verifying 4 to 5 drafted tokens costs roughly the same GPU overhead as checking a single token.
## BACKGROUND
Speculative decoding accelerates LLM inference by using a small draft model to generate candidate tokens that a larger target model verifies in parallel, replacing slow sequential token generation. A CUDA megakernel is a GPU optimization technique where multiple distinct computational steps are fused into a single large kernel, avoiding the overhead of repeatedly calling the GPU from the CPU.