Sakana AI Launches Fugu Cyber, Outperforming Frontier Models in Cybersecurity Benchmarks
Japanese AI startup Sakana AI has launched Fugu Cyber, a multi-agent cybersecurity model designed to handle complex modern cyber defense tasks. The model achieved an 86.9% success rate on the CyberGym benchmark and 72.1% on CTI-REALM, outperforming frontier models like OpenAI's GPT-5.5-Cyber and Anthropic's Claude Mythos-Preview. This release demonstrates the power of multi-agent orchestration systems in specialized domains like cybersecurity, where complex, multi-step workflows are common. By outperforming frontier models, Fugu Cyber highlights a shift toward task-specific, agentic architectures over general-purpose LLMs for enterprise security. Fugu Cyber operates as a single API that dynamically coordinates specialized agents to execute multi-step tasks, simplifying integration for developers. In addition to the API, Sakana AI provides on-premise deployment services supported by their enterprise application team.
## BACKGROUND
Sakana AI's Fugu framework is a multi-agent orchestration system packaged as a single model API, allowing users to access specialized agents through an OpenAI-compatible interface. The benchmarks used to test Fugu Cyber include CyberGym, which evaluates real-world vulnerability analysis, and CTI-REALM, a Microsoft-developed benchmark for evaluating threat intelligence detection engineering.