Developer Tests Gemma4-31B and Qwen3.8-27B for Real-World Coding Tasks
A software engineer evaluated local open-weights models Gemma4-31B and Qwen3.8-27B under 4-bit Unsloth quantization alongside proprietary GPT-6.1-Sol on repository analysis and C# app creation. Gemma4 demonstrated higher execution efficiency and focus, whereas Qwen3.8 proved more persistent and delivered higher overall code quality after revisions. This evaluation highlights the pragmatic trade-offs between speed, persistence, and output quality when deploying open-weight models locally on consumer hardware for software engineering. It also underlines the remaining gap between local 30B-class models and top-tier proprietary frontier models in complex code generation. During codebase analysis, Gemma4 required only ~10 server calls compared to Qwen3.8's 20-30 calls while achieving equivalent analytical depth. For multi-file project authoring, Gemma4 yielded a B- result needing multiple feedback passes, Qwen3.8 reached a viable B+ after one revision, and GPT-6.1-Sol delivered an A+ one-shot solution.
## BACKGROUND
Open-weight Large Language Models (LLMs) can run directly on consumer hardware, enabling offline code generation without API fees or privacy risks. Quantization methods like Unsloth 4-bit compress model weights into 4-bit representations, significantly reducing VRAM footprint so 27B-31B models can run efficiently on desktop GPUs.