
OpenAI's AI Scorecard: Measuring What Really Matters
OpenAI CFO Sarah Friar introduces a practical AI scorecard to measure real-world ROI through useful work, cost per successful task, dependability, and return on compute.
5 stories tagged Metrics

OpenAI CFO Sarah Friar introduces a practical AI scorecard to measure real-world ROI through useful work, cost per successful task, dependability, and return on compute.

Researchers created a way to measure how well AI tools work together. It helps AI assistants use the right tools at the right time, making them faster and more reliable.

A new study shows that common methods for evaluating AI error detection can be misleading. The research introduces a controlled stress-test protocol called ErrorBench to reveal these flaws.

A large-scale study found that AI judges often overstate their accuracy, relying on flawed metrics that don't correct for chance agreement. The research evaluated 21 judges across 118 runs and over 541,000 judgments, revealing significant issues with reliability and bias.

Researchers propose new metrics to evaluate AI systems in rule-governed environments, addressing flaws in traditional agreement-based evaluation methods. The Defensibility Index and Ambiguity Index aim to better assess AI decision-making stability and policy compliance.