Seeking an Open-Weights Coding Model to Balance Speed and Performance on RTX 5090
A local LLM user on r/LocalLLaMA asked for open-weights coding model recommendations that bridge the performance and speed gap between Qwen 3.8 27B and Qwen 3.8 Flash Next. Running on an Nvidia RTX 5090 GPU with 96GB of RAM, the user is targeting a decode speed of 75–100 tokens per second for fast iterative coding. As high-end consumer GPUs like the RTX 5090 become accessible, developers running models locally must continually fine-tune their setups to balance intelligence against generation speed. Finding an ideal middle ground or combining fast models with smarter ones reflects a common workflow challenge in the local AI coding ecosystem. The user currently gets over 200 tokens per second (TPS) on Qwen 3.8 27B but finds its code quality insufficient, whereas Flash Next delivers higher quality but runs at a slower 50 TPS. To hit their 75–100 TPS target, the user is considering upgrading memory to 128GB DDR5 or using Flash Next for high-level planning while reserving 27B for code execution.
## BACKGROUND
Qwen is a family of open-weights language models developed by Alibaba Cloud, encompassing dense models as well as Mixture-of-Experts (MoE) architectures tailored for reasoning and coding. Local inference speed, measured in tokens per second, depends heavily on GPU memory bandwidth, model size, and memory offloading between VRAM and system RAM.