~/LLM INFERENC/strata-engine-runs-125b-qwen3-8-moe-model-on-12gb-consumer-gpus

Strata Engine Runs 125B Qwen3.8 MoE Model on 12GB Consumer GPUs

Developer Niko1221 has open-sourced Strata, an inference engine capable of running quantized 125B parameter MoE models such as Qwen3.8-Flash-Next on consumer GPUs with 12GB of VRAM. By combining expert offloading with speculative decoding, Strata achieves generation speeds of up to 94 tokens per second on an NVIDIA RTX 5070. Strata drastically lowers the hardware requirements for running massive open MoE models locally without relying on expensive enterprise server GPUs. This advances accessible local AI deployment by delivering real-time interactive performance on mainstream desktop setups. The engine stores the entire MoE model in system RAM while caching frequently accessed expert weights in GPU VRAM, using a smaller draft model to accelerate output via speculative decoding. Benchmarks show an RTX 5070 (12GB VRAM) with 64GB RAM generating 94 tokens/sec on Q2_0 quantization and 53 tokens/sec on IQ3_S, with minimum requirements set at 12GB VRAM, 32GB system RAM, and 80GB storage.

## BACKGROUND

Mixture-of-Experts (MoE) models partition neural networks into specialized sub-networks ('experts'), routing inputs to only a small subset of parameters per token. Expert offloading leverages this structure by keeping inactive experts in system RAM and loading active ones into VRAM on demand. Speculative decoding speeds up inference by having a lightweight draft model quickly generate candidate tokens, which the larger target model verifies in parallel without sacrificing accuracy.

## REFERENCES

## KEYWORDS

#LLM Inference#MoE#Local AI#Model Offloading#Speculative Decoding

$ subscribe --daily

Strata Engine Runs 125B Qwen3.8 MoE Model on 12GB Consumer GPUs | Daily News