~/AI SAFETY/chinese-open-source-ai-solution-dognavy-ranks-third-globally-in-cybergym-benchmark

Chinese Open-Source AI Solution DoGNAVY Ranks Third Globally in CyberGym Benchmark

The Chinese AI safety solution DoGNAVY, co-developed by DarkNavy, ranked third globally and first among open-source models in the CyberGym AI safety benchmark with a 90.8% pass rate. It achieved this milestone using only a single open-source model, GLM-5.2, whereas higher-ranking competitors relied on ensembles of closed-source models. This achievement demonstrates that open-source models can compete with proprietary, closed-source ensembles in complex cybersecurity tasks. It lowers the barrier to entry for advanced AI safety capabilities, allowing more innovators to deploy top-tier security agents locally. While Microsoft and Google routed tasks dynamically across multiple top-tier closed-source models, DoGNAVY relied entirely on GLM-5.2, a model designed for long-horizon tasks and coding. CyberGym evaluates AI agents across a dynamic, real-world lifecycle of vulnerability discovery, proof, and patching.

## BACKGROUND

CyberGym is a large-scale, end-to-end cybersecurity benchmark that evaluates AI agents on real-world vulnerability analysis across hundreds of open-source projects. GLM-5.2 is an open-source flagship model released under the MIT license, specifically optimized for long-horizon agentic tasks and coding with a 1-million-token context window.

## REFERENCES

## KEYWORDS

#AI Safety#Cybersecurity#Large Language Models#Benchmarks

$ subscribe --daily

Chinese Open-Source AI Solution DoGNAVY Ranks Third Globally in CyberGym Benchmark | Daily News