Independent Developer Benchmarks Local and API LLMs on Custom Cybersecurity Tasks
An independent developer created Rangebench, a custom 19-task Capture The Flag (CTF) benchmark that evaluates the autonomous hacking capabilities of LLMs operating in isolated Docker environments. Test results showed that the local open-weight Qwen3.8 27B model achieved a 28.1% success rate on its first attempt, whereas commercial API models like GPT-6 Luna solved up to 90.9% of tasks. Amid growing concerns about threat actors leveraging local AI models for cyberattacks, this benchmark empirically demonstrates that mid-sized open models currently pose limited threat on complex exploitation tasks like pwn compared to frontier API models. It provides the open-source AI community with a dynamic framework to evaluate agentic capabilities inside real interactive environments. The evaluation ran local models using llama.cpp RPC distributed inference across two Proxmox nodes equipped with RTX 3090 and 3080 GPUs over a 2.5Gbps link. Across 544 scored attempts, memory corruption and binary exploitation (pwn) proved to be the hardest category for all models, with the highest score reaching only 56%.
## BACKGROUND
Capture The Flag (CTF) benchmarks evaluate AI agents by requiring them to interact with command-line interfaces inside isolated software environments to discover specific secret strings called flags across disciplines like web security, cryptography, and reverse engineering. Distributing LLM inference using protocols like llama.cpp RPC enables offloading matrix computations across multiple remote GPU or CPU nodes, making it possible to host larger models locally across consumer hardware.