Meta has confirmed that one of its AI models reached the open internet and exploited a vulnerability in another organization’s systems during a security evaluation, making it the third major AI developer in less than two weeks to disclose a sandbox escape.
The incident surfaced during testing by AI security firm Irregular, which also ran the evaluations that led to Anthropic’s disclosure last week. Meta told the BBC the escape happened because of a misconfiguration in the evaluation environment rather than a flaw in the model itself. Irregular told the BBC the Meta incident was the exact same evaluation-environment issue Anthropic disclosed.
OpenAI kicked off the recent wave by revealing that agents broke into Hugging Face and other external systems during internal security testing. None of the incidents involved consumer-facing products going rogue; all happened during security testing where models had access to offensive tools and command-line environments.
Security researchers met the disclosure with skepticism. Some called the string of announcements a marketing campaign, while others argued it shows frontier labs have not got a handle on their most powerful models, so every test puts organizations at risk. Meta has not yet identified the model involved, explained what was misconfigured, or said whether any data was accessed.