
Why Auditing AI Agents Means Checking Their Data Sources Too
A new study suggests that to properly audit AI agents, we must also examine the data they use to make decisions. This could change how we ensure AI systems are fair and reliable.
936 stories tagged Research · page 14 of 39

A new study suggests that to properly audit AI agents, we must also examine the data they use to make decisions. This could change how we ensure AI systems are fair and reliable.

Researchers introduced SproutRAG, a new AI framework that enhances how systems understand long documents. It balances detail and coherence better than existing methods, potentially improving tools like document analysis and legal research.

Researchers have developed a method called activation steering to improve synthetic data generation for languages with limited resources. This approach could make AI models more effective and affordable for less common languages.

Scientists identified a critical gap between what AI models intend to do and what they actually execute. This 'intent-execution' gap explains why even advanced AI systems sometimes underperform in real-world tasks.

Researchers have developed a machine-learned comorbidity index that better predicts patient outcomes. Unlike older methods, it captures complex relationships between diseases and health risks.

Researchers have developed a local AI system that can identify and redact student names in classroom transcripts without sending data to third parties. This could make it easier for schools to share educational dialogue for research while protecting privacy.

Researchers have developed a method to help AI models retain detailed audio information, like tone and emotion, while processing speech. This could make voice assistants and transcription tools much more accurate and expressive.

Researchers used an OpenAI reasoning model to help diagnose rare diseases, identifying 18 new diagnoses in previously unsolved cases. This could revolutionize how doctors approach these complex conditions.

A new research paper introduces the concept of distributed general-purpose agent networks where AI agents can collaborate across personal devices, edge nodes, and autonomous computing environments. This could enable more powerful, flexible AI assistants that work together to solve complex problems by sharing data, tools, and permissions. The paper outlines an open peer-to-peer architecture that allows heterogeneous agents to discover and interact with each other, overcoming the limitations of single-agent systems.

OpenAI has introduced LifeSciBench, a benchmark to test AI systems on real-world life science tasks. It's designed and reviewed by experts to ensure AI can handle complex research decisions accurately.

Researchers have identified a critical flaw in AI reasoning: models often reach the same answer through inconsistent paths. They propose a new way to measure this inconsistency to improve AI reliability.

Researchers found that standard parallel sampling in AI search often leads to repetitive queries. Their new method, DivInit, improves efficiency by diversifying initial queries, reducing redundant work.

Researchers created SpeechDx, a comprehensive benchmark for AI to analyze speech for health conditions. This tool could lead to earlier, more accurate medical diagnoses through simple voice recordings.

Researchers introduced MemTrace, a new benchmark that evaluates AI long-term memory by tracking individual knowledge points across three controlled dimensions, rather than aggregating accuracy over independent question rows.

Google's AMIE AI system can now manage complex health conditions as well as primary care physicians. This breakthrough could make high-quality medical advice more accessible to everyone.

Researchers created a benchmark to test if AI models can handle complex executive decisions. The study simulates real-world leadership challenges, like balancing conflicting advice and managing resources under constraints.

A new study tests whether AI language models can discover fundamental mathematical concepts like zero without explicit training, exploring their ability to generalize beyond training data and hypothesize genuinely new mathematical structures.

Researchers developed an AI system that creates digital twins of patients to simulate treatments and optimize care in real-time while adhering to safety constraints. This approach could lead to more personalized and effective medical decisions.

Researchers introduced a self-evolving AI agent that improves legal case searches by rewriting queries to boost BM25 performance without any parameter training. This could streamline legal research for professionals.

Researchers have proposed a new way to define good AI explanations, focusing on counterfactuals and user beliefs. This could help AI systems explain their decisions more clearly to everyday users.

Researchers have developed a new method called Telegraph English that compresses information for AI models while preserving meaning. This could make AI answers more accurate and efficient, especially for complex, multi-step questions.

Researchers have identified memory traces in AI models that behave similarly to biological memory units in human brains. This discovery could help us understand how AI learns and remembers, potentially improving future AI systems.

Scientists have created a method to measure how much AI agents trust each other, using a survival game where verification is costly. This could help design better teams of AI agents for complex tasks.

Researchers introduced PrologMCP, an open-source server that helps AI models solve complex problems by delegating deductive reasoning to Prolog. This could make AI systems more reliable for tasks requiring deep logical analysis.