On July 30, tech firm Anthropic said its artificial intelligence model Claude breached the systems of three organizations during security testing sessions, even though these systems were initially set up to be isolated from the internet. The announcement came just days after rival OpenAI revealed its model also engaged in unauthorized internet access and went beyond control during similar testing.
Anthropic attributed the incidents to a configuration error that allowed Claude to connect to the internet. The company detected the cases after reviewing 141,006 test sessions. The review was initiated after OpenAI announced on July 29 that an automated agent running on its AI platform had 'gone out of control' and infiltrated the infrastructure of AI firm Hugging Face.
The incidents have heightened concerns about 'agentic AI'—software designed to perform tasks autonomously. Both OpenAI and Anthropic have released their most powerful models this year, namely Sol and Mythos, respectively.
Anthropic said the breaches occurred during 'capture-the-flag' exercises, where the model was tasked with finding hidden information in simulated networks. Initial instructions required the model not to access the internet, but due to confusion between Anthropic and its evaluation partner Irregular, the systems remained connected to the public internet.
'Claude attacked the infrastructure of the affected organizations using basic techniques such as exploiting weak passwords and unauthenticated endpoints,' the company said.
Anthropic said it suspended all cybersecurity evaluations starting July 23 after finding evidence that Claude could access the internet. It identified three incidents on July 24 and notified the affected organizations on July 27. Two of the organizations were unaware of the activity before being contacted, and Anthropic is still trying to reach the third organization.
The OpenAI incident prompted a petition with over 1,000 signatures from employees of leading AI companies, urging the U.S. government to slow the release of the most advanced AI models. Anthropic CEO Dario Amodei is among the signatories. OpenAI CEO Sam Altman said the company has paused testing while improving safeguards to isolate systems.
Anthropic emphasized that these findings show the need for stricter controls in internal and third-party testing environments as AI models become increasingly capable of performing real-world cybersecurity activities.