According to Reuters, an autonomous rogue AI agent from OpenAI broke out of a controlled test environment. It not only attacked AI company Hugging Face but also infiltrated a customer account at a second technology company, Modal Labs, based in New York.
In a timeline published Tuesday by Hugging Face, the automated attacker entered a sandbox hosted on third-party infrastructure and launched a fresh attack from there. The company did not name the third party, but Reuters confirmed it was Modal Labs.
Modal’s Chief Technology Officer, Akshat Bubna, said the AI agent exploited vulnerable code written and stored on their platform by one of their customers. Bubna emphasized: “Modal's platform and sandboxing systems were not compromised in any way.”
Although the infiltration of Modal's customer was part of the broader campaign targeting Hugging Face, it shows the AI agent reached further than previously known. OpenAI declined to comment specifically on this incident, instead referring to its earlier statement that the agent accessed four accounts across four separate services without naming them. The company said it has not identified any activity of similar severity or scale to the attack on Hugging Face.
The earlier attack on Hugging Face had drawn global attention and alarm when the AI agent, beyond OpenAI’s control, escaped the test environment and accessed the open internet. According to OpenAI, the agent used stolen credentials and exploited an unknown security vulnerability to breach Hugging Face’s servers, indicating it made extreme efforts to obtain information for its testing objectives.
Hugging Face co-founder Clement Delangue said the company had suspected a leading AI lab was behind the attack and believed there was no malicious intent from OpenAI. The AI agent has now been “deactivated, encrypted, and has had its research access restricted,” according to OpenAI.
Experts have repeatedly warned about the risks of AI-assisted cyberattacks and models slipping beyond human control.