AI Models Cross Capability Threshold in Binary Exploitation Benchmarks
Anthropic's Frontier Red Team reported that frontier AI models including GLM-5.3 and Claude Mythos Preview achieved full control flow hijacks in binary exploitation benchmarks. This marks a notable capability shift, as previous generation models such as Claude Opus 4.6 and GLM-5.2 failed to achieve any successful hijacks on the same benchmark. This milestone demonstrates that advanced LLMs are acquiring non-trivial autonomous offensive cyber capabilities at low levels of software execution. It highlights the growing urgency of proactive AI safety testing and red teaming to assess risks before deploying powerful frontier models. Across 100 internal binary exploitation benchmark tasks, Claude Mythos Preview achieved full control flow hijacks in 6% of trials, while GLM-5.3 succeeded in 4%. Earlier models failed in 100% of trials, indicating a clear qualitative leap in vulnerability exploitation capability.
## BACKGROUND
Binary exploitation involves abusing vulnerabilities in compiled software memory to corrupt execution state and elevate privileges. A control flow hijack allows an attacker to alter the intended path of a program's code execution to run arbitrary commands. AI red teaming adapts security penetration testing methods to evaluate potential misuse and unsafe capabilities in machine learning models.