High-Speed LLM Inference on Consumer Hardware via Custom Strata Engine Fork
Developer bodhi371 demonstrated running the Qwen3.8-Flash-Next model at IQ3_S quantization using a custom experimental fork of the Strata inference engine on a machine with 12GB VRAM, 32GB RAM, and an NVMe SSD. The setup achieved decoding speeds of 20–30 tokens per second and context prefilling speeds reaching up to 90,000 tokens per second at 131k context. This setup demonstrates how aggressive quantization combined with intelligent memory offloading enables running massive Mixture-of-Experts models locally on mid-range hardware. It lowers the barrier to entry for users seeking high-performance local AI without relying on expensive cloud infrastructure or high-end enterprise GPUs. The project relies on architectural modifications to Strata that dynamically push expert layers across VRAM, system RAM, and NVMe storage. Lowering the quantization level to Q2 increased decoding performance further to 39–45 tokens per second.
## BACKGROUND
Strata is an open-source local LLM inference engine designed to run large Mixture-of-Experts (MoE) models on consumer hardware by dynamically offloading dormant expert layers between GPU VRAM, system memory, and high-speed NVMe storage. IQ3_S is an importance-matrix-based quantization technique (i-quant) in the ggml ecosystem that compresses model weights down to approximately 3 bits per parameter while preserving accuracy. Abliterated models are modified LLMs engineered to reduce safety refusal behaviors while keeping general reasoning capabilities intact.