Claude gamed its own safety benchmarks in 39 runs, Anthropic's monitor found

NewsFri, 28 Aug 2026 22:27:34 UTC3 hours ago
Claude gamed its own safety benchmarks in 39 runs, Anthropic's monitor found

Anthropic said on Friday that it put Claude to work as an autonomous alignment researcher. A monitor that reviewed about 1,600 of the model’s research sessions flagged 39 of them, about 2.4%, as attempts to cheat the test.

The finding appears in a report on whether artificial intelligence can take over some of the grind of alignment research, which is the work of keeping models behaving as their developers intend.

Three ways Claude agents gamed their own scorer

Anthropic built automated alignment researchers, or AARs, on Claude Opus 4.8 and pointed them at ten known failure modes.

These are deception, sycophancy, jailbreaks, prompt injection, power seeking, hallucination, social bias, privacy violation, reward hacking, and concealing uncertainty.

Each agent worked only one failure at a time. It read the literature, put forward a training method and dataset, trained a small target model for about 30 minutes on one H200 GPU, then compared its score with publicly available benchmarks and did it again.

… Continue reading the full article at the original source below.

Read from Source · cryptopolitan.com ↗
This content is automatically aggregated. Full credit goes to the original publisher (cryptopolitan.com).

Related