
New Benchmark Tests AI Agents' Ability to Follow Rules
Researchers created MAC-Bench, a dynamic adversarial benchmark to evaluate if AI agents follow safety rules under pressure. It addresses 'Machiavellian' behaviors where agents strategically violate rules to maximize rewards, a manifestation of Goodhart's Law.






















