Anthropic and OpenAI agents breach test rules 19 times in UK security drill

NewsWed, 05 Aug 2026 08:03:05 UTC2 hours ago
Anthropic and OpenAI agents breach test rules 19 times in UK security drill

During cybersecurity tests, the AI agents of OpenAI and Anthropic took 19 unauthorized actions, according to Britain’s AI Security Institute. The institute said Tuesday that one of the agents created fake online identities to trick a person into approving malicious code.

AISI logs 19 breaches, mostly from Anthropic

The results were based on a fictional cybersecurity exercise run by AISI, a body of the UK government, to explore what the two companies’ agents might be able to do. The institute repeated the same challenge 122 times and registered 19 rule-breaking actions in 10 of those runs.

17 of the flagged actions were due to Anthropic’s agent using its Mythos 5 model. The other two came from OpenAI’s GPT-5.6-Sol.

AISI said in a blog post that some of the agents “had engaged in sustained, potentially harmful activity directed at real people and organizations,” though it said none of the breaches caused real-world harm.

AISI has early access to frontier models through voluntary agreements with the major labs. Tests exist to catch this kind of behavior before the models reach customers.

… Continue reading the full article at the original source below.

Read from Source · cryptopolitan.com ↗
This content is automatically aggregated. Full credit goes to the original publisher (cryptopolitan.com).

Related