
New Benchmarks for Measuring AI's Reliability in Healthcare
Researchers propose new ways to test AI models in healthcare to ensure they're safe and reliable. This could make AI tools more trustworthy for doctors and patients.
35 stories tagged Reliability · page 2 of 2

Researchers propose new ways to test AI models in healthcare to ensure they're safe and reliable. This could make AI tools more trustworthy for doctors and patients.

Researchers have developed a way to make AI models better at admitting when they don't know something. This could make AI assistants more reliable in everyday use. The method works without needing to see the model's internal workings, making it useful for commercial AI services.

Researchers have identified a flaw in AI repair systems where rankings change unpredictably. They've released a tool to help developers spot and fix these issues. This could make AI systems more reliable for everyday users.

OpenAI says its latest ChatGPT model, GPT-5.5 Instant, makes up facts 52.5% less often. This could make AI assistants more reliable for everyday use.

Testing AI agents is tricky because they often produce different answers. Researchers are developing new methods to evaluate their performance consistently. This matters for everyone who relies on AI tools for daily tasks.

TrainForgeTester is a new open-source tool that helps you test AI agents in real-world scenarios. It focuses on catching mistakes like wrong tool calls or skipped steps, making AI agents more reliable for everyday use.

A new tool lets users explore the reliability of large language models through interactive data visualizations. It highlights inconsistencies in model responses across different queries.

Operating 14 AI agents for half a year revealed critical insights on scalability, cost, and reliability. The experience highlights the need for robust infrastructure and continuous monitoring.

SymptomWise introduces a hybrid framework that separates language understanding from diagnostic reasoning to eliminate hallucinations in AI symptom analysis. By combining expert-curated knowledge with deterministic inference, the system ensures traceable and consistent outputs in safety-critical settings.

ProofSketcher combines LLMs with a lightweight proof checker for reliable math and logic reasoning. It aims to address the limitations of LLMs in producing persuasive but flawed arguments.
Gemini 3.1 Flash Live is now available, improving audio AI. This update enhances natural and reliable audio processing across Google products.