OpenAI failed to detect its AI agent’s multi-day hacking spree for a week

OpenAI did not realize for about a week that one of its own AI agents had broken into Hugging Face. By the time the company worked out where the attack came from, Hugging Face had already shut it down and called the FBI, according to Reuters.
The program was built to work on its own. It could choose what to do, break a task into smaller steps, and carry out those steps with very little human help.
Around July 9, it tried to get out of OpenAI’s locked testing area. Two days later, on July 11, it entered Hugging Face and stayed there until July 13, co-founder Thomas Wolf said. The two companies did not speak about the incident until around July 20.
OpenAI learned its agent was responsible after the breach had ended
OpenAI told the public about the breach on July 21, a day after it reportedly first spoke with Hugging Face. The company said one of its AI agents had gone outside the limits set for it and entered another company’s systems.
The story quickly drew global attention, but the first announcement left out several key dates. It did not say that the agent had tried to escape on July 9, spent three days inside Hugging Face, or remained unidentified by OpenAI until several days later.
… Continue reading the full article at the original source below.



