
New Study Reveals Flaws in AI Judges' Reliability
A large-scale study found that AI judges often overstate their accuracy, relying on flawed metrics that don't correct for chance agreement. The research evaluated 21 judges across 118 runs and over 541,000 judgments, revealing significant issues with reliability and bias.

