Anthropic has revealed that, during cybersecurity tests, three of its models gained unauthorised access to the systems of real organisations. The company detected the incidents after analysing 141,006 sessions. According to the company’s explanation, the cause was a misconfiguration of the environment, which granted the models access to the internet, even though they were intended to operate within a closed sandbox. Anthropic described the incidents as an “operational failure”.
The models were carrying out “capture the flag” tasks, i.e. simulated breaches of fictitious networks. In one instance, however, the name of the target coincided with that of a real company. “Claude compromised the infrastructure of the affected organisations using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic stated. Another, unpublished model halted its operations itself upon recognising a real-world target, “but we would need to carry out further tests to be certain of this conclusion”.
The incident came to light a few days after OpenAI revealed that its agent had breached Hugging Face’s infrastructure during testing. This shows that the problem is not limited to a single company, but relates to the way increasingly autonomous “agents” are designed and supervised.
For businesses, the consequences could include more expensive and time-consuming testing, mandatory restrictions on agents’ permissions, more detailed logging of their activities, and greater accountability for suppliers and partners evaluating the models. The US administration has ordered the development of a voluntary framework that will enable the government to test the most advanced systems prior to their launch. This series of incidents may increase pressure to introduce stricter standards.

