
New Method 'Metric Match' Reduces Reliance on Human Annotations for AI Judge Evaluation
Researchers developed Metric Match, a subset selection method that accurately estimates LLM judge reliability from limited human annotations, potentially reducing the cost of AI evaluation.
