Ranking Every Language Model That Can Run on 10-16GB VRAM
A comprehensive evaluation and ranking of various open-source language models optimized to run locally on consumer hardware with 10-16GB of VRAM was published, detailing their speed, coherency, and behavioral traits. This ranking provides practical guidance for local AI enthusiasts and developers looking to deploy capable language models within strict consumer hardware constraints, helping them select the best model for specific conversational and roleplay tasks. The evaluations prioritize models achieving at least 10 tokens per second, testing them across metrics like instruction following, dialogue quality, and slop patterns, while ignoring chain-of-thought models.
## BACKGROUND
Running large language models locally typically requires high-end hardware, but recent architectural improvements and quantization techniques allow consumer GPUs with 10-16GB of VRAM to efficiently execute capable models with extended context windows using specialized software like Kobold or Llama.