Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities
Summarized by AI from reporting by ArXiv cs.AI, published under our editorial policy.
Researchers from UC Berkeley and Stanford published a method to audit black-box LLM agents using a low-cost, open-weight surrogate model. The surrogate reads the same context and proposed action as the primary agent and provides a reliable error signal from its log-probabilities, catching mistakes before they execute.

Key takeaways
- Researchers from UC Berkeley and Stanford published a method to audit black-box LLM agents using a low-cost, open-weight surrogate model.
- The surrogate model reads the same context, schema, and proposed action as the primary AI agent and provides a reliable error signal from its log-probabilities.
- Frontier chat APIs hide token probabilities, and agents' stated confidence is barely better than chance on critical mistakes, making current auditing methods unreliable.
- The method catches mistakes before they execute, improving the reliability of AI agents in real-world applications.
Researchers from the University of California, Berkeley, and Stanford University published a paper on ArXiv titled 'Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities.' The paper introduces a method to detect errors in AI agents before they execute actions, potentially preventing costly mistakes.
Why Frontier AI Agents Are Hard to Audit
AI agents, which are systems that use large language models (LLMs) to perform tasks, often make mistakes. These errors can go unnoticed until the agent has already taken action, leading to potential problems. Current methods for detecting these errors are unreliable: frontier chat APIs hide the model's token probabilities, the agents' stated confidence levels are barely better than chance on the mistakes that matter, and resampling does not help because frontier models are highly repetitive, reproducing the same call across samples.
How the Surrogate Model Auditing Works
The researchers propose using a low-cost, open-weight surrogate model to audit the work of the primary AI agent. This surrogate model reads the same context, schema, and proposed action as the agent and provides a more reliable signal of potential errors by recovering the missing signal from the surrogate's log-probabilities. By running the surrogate model in parallel, the system can catch mistakes before they are executed, improving the overall reliability of the AI agent.
Why This Matters for Everyday Users
This research is significant because it addresses a critical issue in AI systems: reliability. As AI agents become more integrated into daily life, from customer service to financial transactions, ensuring they make fewer mistakes is crucial. This method could lead to more trustworthy AI systems, reducing the risk of errors that could have real-world consequences.
What You Can Do Today
While this research is still in the early stages, you can stay informed about advancements in AI reliability. Keep an eye on updates from the researchers and consider using AI systems that implement similar error-detection methods in the future. If you are developing or using AI agents, consider integrating surrogate models to improve their accuracy and reliability.
Frequently asked
- What is a surrogate model in this context?
- A surrogate model is a low-cost, open-weight model used to audit the work of a primary AI agent. It reads the same context, schema, and proposed action as the agent and provides a reliable error signal from its log-probabilities.
- Why can't we just ask the AI agent how confident it is?
- The paper states that frontier chat APIs hide the model's token probabilities, and the agent's stated confidence is barely better than chance on the mistakes that matter, making it an unreliable signal.
- Is this method currently available for use?
- The source paper does not specify a timeline for availability. It is a research publication on ArXiv, so it is still in the early stages.