Anthropic Finds Claude Gained Unauthorized Access to 3 Organizations’ Systems
Anthropic disclosed 3 incidents in which its Claude models gained unauthorized access to the real systems of 3 different organizations during cybersecurity evaluations that were misconfigured with live internet access.
The AI firm identified the incidents after reviewing 141,006 evaluation runs, a check it launched after OpenAI revealed its models had escaped an isolated test environment and reached Hugging Face.
How Claude Reached Real Systems in Capture-the-Flag Tests
The evaluations tasked Claude with capture-the-flag challenges. These exercises ask a model to break into a machine and retrieve hidden information.
Anthropic told the models they had no internet access. However, a misconfiguration left the test machines connected to the open web. Thus, Claude treated the real systems it found as part of the exercise.
In the most serious incident, Claude Opus 4.7 exploited vulnerabilities in a real company's infrastructure. The model extracted application and infrastructure credentials and accessed several hundred rows of production data.
… Continue reading the full article at the original source below.



