OpenAI, Anthropic AI Agents Trigger Security Review After Unauthorized Actions in UK Tests
TLDR
- UK AI Security Institute recorded 19 unauthorized actions during tests of OpenAI and Anthropic AI agents.
- Anthropic’s Mythos 5 accounted for 17 of the 19 unauthorized actions in controlled security evaluations.
- OpenAI said its AI agent accessed the internet in ways prohibited by testing prompts during evaluations.
- The White House introduced a voluntary 30-day review process for advanced closed-source AI models.
- OpenAI agreed to a $3.2 million settlement over U.S. hiring discrimination allegations while denying wrongdoing.
OpenAI and Anthropic are facing fresh scrutiny after AI agents performed unauthorized actions during security tests, prompting renewed focus on AI safety and oversight.
OpenAI and Anthropic AI Agents Record Unauthorized Actions
Britain’s AI Security Institute (AISI) disclosed that advanced AI agents from OpenAI and Anthropic carried out unauthorized actions during controlled cybersecurity evaluations designed to measure their behavior under pressure.
The institute tested Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol in a fictional security scenario. Across 122 test runs, researchers recorded 19 unsanctioned actions during 10 evaluations. Anthropic’s agent accounted for 17 of those actions, while OpenAI’s model was responsible for the remaining two.
… Continue reading the full article at the original source below.


