Running 27B LLM at 262K Context and ~130 tok/s on a Single RTX 4090
A developer created a custom converter and integration for the C++/CUDA inference engine NInfer, allowing the uncensored Qwen3.8-27B model to run on a single RTX 4090 GPU. The optimized setup achieves generation speeds of ~130 tokens/second across its full native 262K context window using Multi-Token Prediction (MTP3). This achievement demonstrates that high-throughput inference for 27B models with ultra-long context windows is viable on high-end consumer hardware. By significantly outperforming general-purpose engines like llama.cpp, it shows the performance potential of dedicated CUDA engines paired with native speculative decoding techniques. The deployment uses MTP3 speculative decoding with a 70.8% acceptance rate and delivers a prompt prefill speed of 3,591 tokens/second on a 9K prompt. Output accuracy remains high with perplexity within 1.3% of the official artifact, while vision multimodal support and tool calling remain fully intact.
## BACKGROUND
NInfer is a specialized C++/CUDA inference engine designed to achieve maximum single-GPU performance on consumer hardware like NVIDIA RTX GPUs. Multi-Token Prediction (MTP) is a form of speculative decoding where an LLM natively predicts multiple future tokens per step, speeding up generation without needing a separate draft model.