LlamAmpere v0.4 Achieves 95+ TPS for 27B LLM on Single RTX 3090
Developer JakeATX released LlamAmpere v0.4, an Ampere-optimized fork of llama.cpp that achieves over 95 tokens per second across a 100,000-token generation for a 27B model on a single RTX 3090 GPU. The update delivers a ~10% speed boost and a 10%+ context length expansion over the previous release, supporting context windows up to 262K tokens. This project demonstrates how target-specific CUDA kernel optimizations combined with KV cache quantization can push consumer-grade hardware to run high-parameter models at extreme context lengths. It enables local AI users to achieve inference speeds that rival or beat enterprise-focused engines like vLLM without requiring high-end datacenter GPUs. The benchmark was conducted using an IQ4_XS-M 4.6 bits-per-weight quant of Qwen 27B with TurboQuant (TQ5/TQ4) KV cache quantization at temperature=1 to reflect realistic generation speed. The author noted that TurboQuant KV quantization introduces less than one-third of the Kullback–Leibler divergence observed when moving from 8-bit to 4.6-bit weights, showing no statistically significant accuracy degradation at the task level.
## BACKGROUND
Llama.cpp is a popular C/C++ open-source inference library optimized for executing Large Language Models locally on consumer hardware using GGUF quantization formats. Running 27-billion-parameter models with context windows up to 262K tokens typically demands massive VRAM; memory-saving techniques like KV cache quantization and Ampere-architecture optimizations make these large context windows possible on 24GB GPUs like the NVIDIA RTX 3090.