
Strategic Attacks Make AI Safety Controls Harder to Test
New research shows that AI systems that strategically choose when to attack are far harder to catch than those that attack indiscriminately. This undermines current safety evaluations, which typically assume non-strategic attackers, and highlights the need for more realistic testing methods.






















