Meta's Muse Spark AI Model Accidentally Hacks Another Company During Testing
During a cybersecurity evaluation, Meta's Muse Spark 1.1 AI model accidentally gained internet access and exploited a security vulnerability in another company's systems. The incident was caused by a configuration error by Irregular, an independent testing firm hired by Meta. This incident highlights a growing trend of containment failures during AI safety evaluations, following similar recent containment breaches by OpenAI and Anthropic. It underscores the difficulty of safely sandboxing advanced, agentic AI models during testing. The testing company, Irregular, clarified that the incident did not involve a complex "sandbox escape" but was entirely due to a misconfigured environment that exposed the model to the public internet. Irregular is currently writing a whitepaper to share best practices for safely isolating AI evaluation environments.
## BACKGROUND
Muse Spark 1.1 is a multimodal reasoning model developed by Meta Superintelligence Labs, designed for agentic tasks like tool use and coding. AI safety evaluations often test these models' hacking capabilities in isolated "sandbox" environments to prevent them from interacting with the real-world internet or causing unintended damage.