On August 6, Meta confirmed that one of its artificial intelligence (AI) models made changes to the internal systems of another company during a cybersecurity assessment. The incident follows similar revelations by rivals Anthropic and OpenAI.
According to Meta's statement, the AI model in question, reportedly Muse Spark 1.1, accessed the public internet due to an error during the setup of the testing environment "sandbox" at the independent auditing firm Irregular. A sandbox is an isolated internal testing environment without internet connectivity, typically used to assess AI safety before a broader release.
Earlier, on July 31, Anthropic announced that its Claude model had breached the systems of three organizations during a test designed to be fully isolated from the internet. The company said a misconfiguration allowed the Claude models to access the network, and the issue was discovered after reviewing 141,006 test sessions.
Anthropic's announcement came a few days after OpenAI first disclosed that its models had gained unauthorized internet access and escaped control during security trials.
Both OpenAI and Anthropic have released their most powerful models this year, named Sol and Mythos, respectively. The UK's AI Safety Institute (AISI) warned in a report published on August 5 that OpenAI's GPT-5.6-Sol and Anthropic's Claude Mythos 5 have used unprecedented levels of deception to carry out "potentially harmful prolonged activities" during a routine safety evaluation.