ISTA-DASLab Releases GSQ and RCO Optimized GGUFs for Qwen3.8-Flash-Next
ISTA-DASLab has released quantized GGUF weights for the 176.9B parameter Qwen3.8-Flash-Next MoE model utilizing new GSQ quantization and RCO optimization techniques. The release features standard quantized GGUFs ranging from 2.40 to 3.50 bpw as well as an expert-pruned Coder variant that removes half of the model's routed experts to reach an effective average bitwidth of ~1.89 bpw. This development enables running a massive 176.9B-parameter model, originally requiring 354 GB at BF16, on a single 32 GB accelerator by bringing the resident memory working set down to 29.6 GB. It proves that combining advanced post-training quantization with targeted expert pruning can drastically reduce memory requirements while retaining over 98% of coding benchmark performance. GSQ uses Gumbel-Softmax relaxation for scalar post-training quantization compatible with standard GGUF formats, while RCO applies Riemannian constrained optimization to enforce precise per-tensor bit allocations and select experts to retain. The 50% pruned Coder variant eliminates 256 of 512 experts per layer while keeping retained weights at 3.5 bpw, scoring 75.60 on SWE-bench Verified (91.3% of the unpruned BF16 model).
## BACKGROUND
Mixture-of-Experts (MoE) LLMs distribute model parameters across specialized sub-networks called experts, activating only a subset per token to maintain high capacity with manageable compute costs. Quantization compresses models by storing weights with fewer bits per parameter, while GGUF is the popular binary file format used by local inference engines like llama.cpp.