Local Qwen3.8 Setup Outperforms Claude Code on Airbench Benchmark
A local AI setup combining a 3.5-bit quantized Qwen3.8-Flash-Next model, the Strata inference engine, and the pi agent harness scored 100% on the Airbench evaluation suite in 11 minutes on an RTX 5090 GPU. This beat Claude Code Opus 5.5, which achieved the same 100% score but took 14 minutes. This demonstrates that optimized open-weight Mixture-of-Experts (MoE) models running locally on consumer hardware can match top proprietary cloud AI models in practical utility tasks with faster execution times. It highlights the rapid maturation of local inference engines, aggressive quantization techniques, and smart memory offloading strategies. The setup utilized ISTA-DASLab's GSQ-RCO IQ3_S quantization of Qwen3.8-Flash-Next, requiring a GPU with 24 GB VRAM alongside 64 GB of system RAM to offload roughly 50 GB of expert parameters. Strata served a 256k context window, while a lightweight proxy script handled token compaction to prevent context overflow during multi-step benchmark prompts.
## BACKGROUND
Mixture-of-Experts (MoE) language models scale overall parameter counts while activating only a sub-network of experts per token to maintain high inference efficiency. Inference engines like Strata allow massive 100B+ MoE models to run on single consumer GPUs by offloading inactive expert weights to host RAM. Advanced quantization methods like GSQ and RCO reduce model memory footprints down to ultra-low bitrates while preserving base performance.