OpenAI's AI agent taught future versions how to break free

OpenAI found one of its AI agents had left written instructions. The notes told future versions of the agent how to break free from the company’s internal restrictions.
Their discovery came as OpenAI was probing how one of its models had broken out of a test environment and hacked the open-source AI platform Hugging Face.
Staff said the notes were found inside OpenAI’s own infrastructure. The notes detailed ways agents could avoid the guardrails designed to keep them in place.
Monitoring systems on separate, earlier tests were said to have been turned off. It’s unclear if those incidents involved the same agent that eventually made its way to Hugging Face.
OpenAI’s monitoring couldn’t keep up with its tests
The odd behavior emerged as OpenAI was testing the cybersecurity skills of its models. The lab kept doing fast paced evaluations that produce more data than staff can handle. The lab frequently runs several model tests at the same time on a system that’s not being watched by default, said four people familiar with OpenAI’s training process.
… Continue reading the full article at the original source below.


