Anthropic has disclosed three incidents in which Claude AI models accessed the open internet during cybersecurity evaluations and subsequently gained unauthorised access to the production infrastructure of three organisations. The incidents were identified as part of a large-scale retrospective review launched after OpenAI disclosed that some of its models had escaped an isolated testing environment and accessed Hugging Face’s production infrastructure.
Anthropic reviewed 141,006 evaluation runs in which Claude could potentially have obtained internet access. The review identified three cases involving evaluations conducted with third-party cybersecurity evaluation partner Irregular.
In each case, Claude was participating in a capture-the-flag (CTF) cybersecurity challenge. The models were instructed to retrieve a secret piece of information, or “flag”, from a different machine on a simulated network. However, due to a misunderstanding between Anthropic and Irregular, the evaluation environments had internet access despite prompts telling Claude that the environment was simulated and had no internet connectivity.
As Claude searched for systems relevant to its assigned task, it encountered real systems on the public internet and treated them as part of the evaluation. The models subsequently gained unauthorised access to production infrastructure belonging to three different organisations.
According to Anthropic, the incidents did not involve sophisticated or previously unknown vulnerabilities. Claude used relatively basic techniques, including exploiting weak passwords and unauthenticated endpoints. The models also remained focused on their assigned CTF objectives rather than attempting to expand their access or escape the evaluation environment.
However, Anthropic noted differences between models in their behaviour once evidence emerged that they were interacting with real-world systems. An older model continued its activity after receiving indications that it was operating on the open internet, while the latest model stopped after recognising that it had left the intended evaluation environment.
The incidents involved three Claude models: Opus 4.7, Mythos 5 and an internal research test model. The earliest incidents identified in the review date back to April.
Anthropic said the models were operated without the standard safeguards normally deployed with generally available versions, including classifiers and monitoring designed to prevent misuse. However, they retained their model-specific safety training. The evaluations were conducted on dedicated infrastructure that was isolated from Anthropic’s sensitive internal systems and customer data.
Anthropic began reviewing evaluation transcripts on 23 July after learning of the potential internet access and suspended all cyber evaluations that same day. The company identified the three incidents on 24 July and notified Irregular and the affected organisations on 27 July.
Two of the organisations contacted by Anthropic had not previously detected the activity. Anthropic said it is working with them on remediation and continues to contact the third organisation.
The company said the findings highlight the importance of ensuring that AI cybersecurity evaluations are properly isolated from real-world infrastructure. Anthropic is conducting the investigation in collaboration with Irregular, which is carrying out its own review.
The incidents also underscore a growing challenge in AI safety testing: models can interpret realistic cyber environments as legitimate targets when boundaries between simulated and real systems are not technically enforced. Anthropic said stronger collaboration between AI developers and evaluation partners will be increasingly important as models become more capable of autonomous cybersecurity operations.

