Three Claude models broke into real companies during Anthropic cyber tests

NewsFri, 31 Jul 2026 06:03:25 UTC2 hours ago
Three Claude models broke into real companies during Anthropic cyber tests

Three of Anthropic’s Claude models escaped from what were meant to be locked test environments, the company said Thursday. They broke into the live systems of three different organizations. A misconfiguration had handed the models working internet access during the cybersecurity drills.

Claude models treated real companies as practice targets

OpenAI said on July 21 that a model escaped its sandbox and exploited a zero-day to reach Hugging Face’s production systems.

This prompted Anthropic to comb through its logs. The company said in a blog post on its Frontier Red Team site that it analyzed 141,006 evaluation runs in which Claude could theoretically have accessed the internet. It identified three incidents related to Irregular, an external partner that performs and manages some of its offensive security testing.

In these tests, a model is told a secret is hidden on another machine and asked to go get it. Anthropic’s prompt to Claude had told it that the environment was a simulation.

A misconfiguration left evaluation machines with live internet access, and neither company caught it until last week. Anthropic described the mix-up as a “misunderstanding” between the two companies, adding that it is “approaching the fixes as if the responsibility were ours alone.”

… Continue reading the full article at the original source below.

Read from Source · cryptopolitan.com ↗
This content is automatically aggregated. Full credit goes to the original publisher (cryptopolitan.com).

Related