~/LOCALLLM/running-27b-llm-at-262k-context-and-130-tok-s-on-a

Running 27B LLM at 262K Context and ~130 tok/s on a Single RTX 4090

A developer created a custom converter and integration for the C++/CUDA inference engine NInfer, allowing the uncensored Qwen3.8-27B model to run on a single RTX 4090 GPU. The optimized setup achieves generation speeds of ~130 tokens/second across its full native 262K context window using Multi-Token Prediction (MTP3). This achievement demonstrates that high-throughput inference for 27B models with ultra-long context windows is viable on high-end consumer hardware. By significantly outperforming general-purpose engines like llama.cpp, it shows the performance potential of dedicated CUDA engines paired with native speculative decoding techniques. The deployment uses MTP3 speculative decoding with a 70.8% acceptance rate and delivers a prompt prefill speed of 3,591 tokens/second on a 9K prompt. Output accuracy remains high with perplexity within 1.3% of the official artifact, while vision multimodal support and tool calling remain fully intact.

## BACKGROUND

NInfer is a specialized C++/CUDA inference engine designed to achieve maximum single-GPU performance on consumer hardware like NVIDIA RTX GPUs. Multi-Token Prediction (MTP) is a form of speculative decoding where an LLM natively predicts multiple future tokens per step, speeding up generation without needing a separate draft model.

## REFERENCES

## KEYWORDS

#LocalLLM#LLM-Inference#CUDA#RTX-4090#Open-Source-AI

$ subscribe --daily

Running 27B LLM at 262K Context and ~130 tok/s on a Single RTX 4090 | Daily News