~/LOCAL LLM/local-llm-tool-use-benchmark-ranks-qwen3-8-27b-first-and-1

Local LLM Tool-Use Benchmark Ranks Qwen3.8-27B First and 1-Bit Bonsai 27B Last

A developer benchmarked 15 local LLMs for agentic tool-use tasks using the deterministic Toolery benchmark across 143 multi-step scenarios. The evaluation placed qwen/qwen3.8-27b in first place with a 71.8% overall score, while the 1-bit quantized prism-ml/bonsai-27b ranked last at 50.5%. As open-weight LLMs are increasingly deployed locally for autonomous agent workflows, empirical benchmarks help developers understand how severe model compression impacts real-world task execution. The results highlight that extreme quantization can cause specific behavioral flaws, such as models failing to recognize when to stop calling tools under budget constraints. The primary failure mode for Bonsai 27B was exceeding tool budget constraints in 103 of its 148 failed trials, showing it knew which tools to use but repeatedly made unnecessary extra calls. Additionally, Bonsai was the slowest model tested with a total runtime of 6,208 seconds, whereas granite-4.2-8b had the fewest overall failures with 69.

## BACKGROUND

Toolery is a deterministic benchmark framework designed to evaluate LLM tool-calling capabilities across 143 hand-written scenarios divided into four difficulty tiers. Bonsai 27B by PrismML is a multimodal 27-billion parameter model quantized down to 1-bit/ternary weights to fit into roughly 4 GB of RAM for execution on consumer hardware.

## REFERENCES

## KEYWORDS

#local-llm#agentic-ai#tool-use#llm-benchmarks#open-source-ai

$ subscribe --daily

Local LLM Tool-Use Benchmark Ranks Qwen3.8-27B First and 1-Bit Bonsai 27B Last | Daily News