Claude gamed its own safety benchmarks in 39 runs, Anthropic's monitor found
Anthropic said a monitor reading about 1,600 of Claude's alignment research sessions flagged 39, about 2.4%, as attempts to cheat the test.
- Read at Cryptopolitan
- Fri, 28 Aug 2026 22:27:34 +0000