Independent Evaluation Compares Prism-ML Bonsai 2 QAT Models on Qwen 3.8
Byteshape published an independent benchmark evaluation of Prism-ML's new Bonsai 2 Quantization-Aware Training (QAT) models based on Qwen 3.8. The evaluation tested the models using Byteshape's standardized Instruct and Thinking benchmark suites to provide a consistent comparison against standard Qwen 3.8 quantizations. Different providers often report benchmark scores using varying evaluation frameworks, runtimes, and sampling parameters, making direct comparisons difficult. This independent analysis gives local LLM developers a reliable side-by-side comparison of Bonsai 2's trade-offs between throughput, model size, and generation quality. Bonsai 2 achieved roughly 91.5% on Byteshape's composite benchmark while maintaining a strong balance between throughput and quality. The tests utilized Prism's specific runtime while keeping all evaluation workloads, reasoning effort settings, and sampling parameters identical to previous Qwen 3.8 runs.
## BACKGROUND
Quantization-Aware Training (QAT) adapts neural network weights during training to better preserve accuracy when models are compressed to low-bit formats like ternary or 1-bit representations. Prism-ML specializes in building extreme low-bit Bonsai models designed to run efficiently on edge devices and local hardware.