Meta's model joins ranks of autonomous cyberattackers
Meta blames now-notorious third party tester for containment failure

Meta is the latest frontier lab to announce that one of its AI models reached the open internet during a cybersecurity evaluation and attacked a third-party service. It is the third company to make such a disclosure in roughly two weeks. Like OpenAI and Anthropic, Meta had no idea what was happening when it happened. (Either that, or they knew and didn’t want to say anything. I’m not sure which is worse.)
The disclosure was short on details: which model was responsible, when the model was loose, how long the model operated unsupervised, or who it attacked. Blame was foisted upon Irregular, the same independent evaluator whose misconfigured tests let Anthropic’s models attack three real companies last week. Now, don’t get me wrong, I understand that Irregular made some errors. At some point, though, you have to turn your gaze to the people who decided to keep building systems they were warned that they couldn’t contain.
CNN reports that Irregular is “developing a white paper to share best practices for containment and securely running cyber evals.” I simply don’t think containment is feasible in the long run. As an anonymous OpenAI employee said, “it’s impossible to patch every single thing that a creative AI can do.”
An anonymous source claims that Meta’s announcement concerns the model Muse Spark 1.1. If that’s true, it’s interesting: Early last month, Irregular claimed that Muse Spark 1.1 did not “materially alter the cyber threat landscape in its current form,” partly because the model was incapable of automating a cyberattack from start to finish. Now, scarcely a month later, we learn that Muse Spark attacked another company’s system and (per Reuters) “altered its internal environment.” It makes me wonder how accurate Irregular’s threat assessments are, across the board.
Because Meta’s Llama series of models is open-weight, I want to clarify that the Muse Spark line is not open-weight (for now, at least). We don’t know very much about the White House’s new AI safety framework, but it plausibly applies to Muse Spark. (With so much still under wraps, it’s hard to say whether any version of Muse Spark will ever qualify as “state-of-the-art.”) That’s bare consolation, though. If the Muse Spark line does get an open-weight model, then people will have free access to a model whose predecessor has, entirely unprompted, engaged in autonomous cyberattacks.
The analyses and opinions expressed on AI StopWatch reflect the views of the individual contributors and the sources they cover, and should not be taken as official positions of the Machine Intelligence Research Institute.
You can receive emails of dispatches as we write them, or subscribe to our Daily Digest for a once-a-day compilation.


