~/AI ML/underdog-ai-releases-saluki-27b-a-2-bit-quantized-7-89gb-agent

Underdog AI Releases Saluki 27B, a 2-Bit Quantized 7.89GB Agent Model

AI research lab Underdog AI released Saluki 27B, a fine-tuned 2-bit quantized version of Qwen3.8-27B compressed to a footprint of just 7.89 GB. Designed specifically for AI agent applications with text-only I/O, the model runs natively using standard llama.cpp. Running a 27-billion-parameter LLM locally typically requires over 20GB of VRAM, but Saluki 27B allows complex models to run on standard consumer GPUs with less than 8GB of VRAM. It also demonstrates that targeted fine-tuning alongside aggressive 2-bit quantization can preserve or even enhance domain-specific agentic skills like tool execution. Saluki 27B outperforms the unquantized Qwen3.8-27B in tool selection and multi-tool parallel function calling, though its mathematical reasoning capabilities are reduced. It uses standard GGUF formatting for llama.cpp, enabling instant compatibility with popular local inference applications without custom builds.

## BACKGROUND

Model quantization shrinks the memory footprint of large language models by converting high-precision weights (such as 16-bit floating points) into lower-bit formats like 2-bit integers. Open-source libraries like llama.cpp facilitate fast, local LLM inference on consumer hardware like desktop GPUs and laptops.

## REFERENCES

## KEYWORDS

#AI/ML#Model Quantization#LLM#Local AI#AI Agents

$ subscribe --daily

Underdog AI Releases Saluki 27B, a 2-Bit Quantized 7.89GB Agent Model | Daily News