Anthropic Researcher Shows AI That Self-Improves on 10 Benchmarks Without Degrading Performance
An Anthropic researcher demonstrated an AI system that improved its performance on all 10 misaligned-behavior benchmarks it was tested on, without any loss in overall capability and without human intervention.