~/LOCAL LLMS/benchmark-shows-harness-choice-impacts-local-llm-performance-as-much-as-model

Benchmark Shows Harness Choice Impacts Local LLM Performance as Much as Model Selection

A developer launched airbench.ai, a crowdsourced benchmark leaderboard evaluating combinations of local LLMs, quantization formats, and execution harnesses across coding, vision, and computer-use tasks. The evaluation shows that optimized local setups can match proprietary cloud models like Claude Code Opus 5.5, with some local configurations completing tasks even faster. The benchmark reveals that the execution harness—the software infrastructure managing prompts, tool calls, and execution loops—can cause a model's success rate to fluctuate from 22% to 96% on identical hardware. This demonstrates that optimizing agent scaffolding is just as critical as model size or architecture when deploying effective local AI agents. Across tests on a single RTX 5090 GPU, `opencode` and `omp` proved to be the most reliable harnesses by maintaining above 85% accuracy across models, while `swift-1.5-qwen3.8-27b-q6_k` was the most robust model. Conversely, Multi-Token Prediction (MTP) variants and expanded 65k context windows consistently degraded performance scores across tested configurations.

## BACKGROUND

An execution harness (or agent scaffolding) is the software framework surrounding a Large Language Model that manages memory, tool execution, and interactive feedback loops. Quantization formats, such as NVFP4 or GGUF IQ variants, compress model weights to allow high-parameter open models to run locally on consumer GPU hardware like NVIDIA's RTX series.

## REFERENCES

## KEYWORDS

#Local LLMs#LLM Benchmarks#AI Agents#Model Quantization#Machine Learning

$ subscribe --daily

Benchmark Shows Harness Choice Impacts Local LLM Performance as Much as Model Selection | Daily News