Anthropic's Claude AI models during internal cybersecurity testing.
Anthropic AI disclosed that several of its Claude models breached the systems of three organisations during internal cybersecurity evaluations after a testing error accidentally granted them access to the public internet. The company said the incidents differed from a recent OpenAI case because the security lapse resulted from human error rather than autonomous AI behaviour.
FBR Collections Tax Revenue Beats July Target by Rs40 Billion
Anthropic reviewed 141,006 cybersecurity test sessions after OpenAI recently reported that one of its AI agents independently exploited a vulnerability during testing. The review uncovered three incidents involving Claude Opus 4.7, Claude Mythos 5 and an internal research model.
According to the company, the models accessed external systems by exploiting weak passwords and unsecured endpoints. Anthropic did not identify the affected organisations but confirmed that the incidents occurred in controlled evaluation environments designed to measure the offensive cyber capabilities of its AI systems.
One incident involved Claude Opus 4.7 targeting a fictional company that shared the name of a real business. The model located vulnerabilities, accessed credentials and reached a database after mistakenly assuming the real-world systems formed part of the test environment.
In another case, an unreleased research model recognised that it had reached a genuine target and voluntarily stopped the attack. Anthropic described that behaviour as encouraging but said it required further testing before drawing broader conclusions.
The company suspended all cybersecurity evaluations on July 23 and informed the affected organisations on July 27. Two organisations were unaware of the activity before Anthropic contacted them, while the company continues to communicate with the third organisation.
Anthropic called the incidents an operational failure and pledged stronger safeguards for both internal and third-party testing environments. It said future evaluations will include tighter controls as AI models become more capable of conducting advanced cyber operations.
The disclosure comes days after OpenAI revealed that one of its AI agents compromised infrastructure belonging to AI platform Hugging Face during separate cybersecurity testing. The incidents have intensified concerns among researchers and policymakers about AI security risks as leading developers race to release increasingly powerful systems.
Industry experts warn that more advanced AI models could become increasingly capable of exploiting vulnerabilities unless developers strengthen oversight, testing procedures and security controls.
