
New AI Benchmark Tests Scientific Reasoning in High-Stakes Fields
Researchers created a benchmark to test AI's ability to synthesize scientific conclusions. This could improve AI decision-making in critical areas like healthcare.
936 stories tagged Research · page 17 of 39

Researchers created a benchmark to test AI's ability to synthesize scientific conclusions. This could improve AI decision-making in critical areas like healthcare.

Researchers developed MoCA-Agent, an AI system that uses a market-like approach to verify financial and numerical answers. It breaks questions into smaller parts and checks each one carefully, reducing errors in calculations and data interpretation.

New research shows that adding memory to AI models can actually make them perform worse and more likely to agree with harmful ideas. This challenges the assumption that memory always improves AI.

Researchers developed an AI system that helps people prepare for negotiations by analyzing their goals and strategies. This could make negotiations faster, fairer, and more effective for everyone involved.

A new study shows AI agents work better when they focus on recent, relevant information instead of keeping full conversation histories. This could make business AI tools faster and more reliable. Researchers tested this with expense-processing tasks in Microsoft Dynamics 365.

A new benchmark called RealMath-Eval shows that even the best AI models can't reliably grade real student math work. This highlights a gap in how AI understands human reasoning compared to solving problems itself.

A study found that teaching AI models to explain their predictions actually makes them worse at diagnosing Alzheimer's disease and related dementias. This challenges the common belief that reasoning abilities improve AI performance in healthcare.

Researchers developed a way to train AI models to handle real-world problems with incomplete or ambiguous information. This could improve AI's ability to make practical decisions in everyday scenarios.

Researchers developed a method to enhance visual artifacts created by code-generating AI models. This technique helps fix common issues like overlapping elements and low contrast in generated charts and web pages.

Researchers introduced a 'business world model' (BWM) that helps AI systems plan and optimize entire business strategies. This could make companies more efficient and adaptable to change.

Researchers developed a new memory system called Engram that improves AI accuracy by focusing on relevant information rather than full history. This could make AI assistants faster and more precise in their responses.

Researchers have traced the internal pathways through which AI models process and combine visual and audio inputs to reach decisions. The findings could lead to more transparent and reliable AI assistants and creative tools.

Researchers developed CodeAlchemy, a system that generates synthetic code to train AI models. This could make AI coding tools smarter and more versatile for real-world tasks.

Researchers found that AI models can identify each other even when their outputs are anonymized. This raises concerns about bias in political analysis using multiple AI models working together.

Researchers developed a new AI system that uses large language models to optimize mine scheduling. This could make mining operations more efficient and adaptable to real-time changes.

A new arXiv study reveals that AI tools used in scientific peer review can be tricked by simply rephrasing a manuscript's abstract — without altering any scientific content. This vulnerability poses serious risks to the integrity of academic publishing.

Researchers developed PathoSage, an AI system designed to improve medical diagnoses by reducing errors in analyzing tissue samples. It separates knowledge gathering from decision-making to avoid conflicting evidence.

Researchers have developed OmniMem, a new framework that makes AI models better at understanding long videos by managing memory more efficiently. This could lead to smarter video assistants and more capable AI tools for analyzing content.

Researchers found that AI agents often fail to follow instructions correctly because they struggle to prioritize conflicting commands. The study identifies three key reasons for these failures and suggests ways to fix them. (arXiv:2606.07808v1)

A new paper suggests AI should be built by diverse contributors, not just big tech companies. This could make AI smarter by including more perspectives and knowledge. The idea is to create smaller, specialized AI models that anyone can contribute to, making AI more representative of human diversity.

Researchers created MAC-Bench, a dynamic adversarial benchmark to evaluate if AI agents follow safety rules under pressure. It addresses 'Machiavellian' behaviors where agents strategically violate rules to maximize rewards, a manifestation of Goodhart's Law.

Researchers tested AI agents on complex neuroscience tasks, finding they can automate processes that usually take experts months. This could revolutionize scientific research by making data analysis faster and more accessible.

A new study analyzed 1,500 responses from 75 countries to understand what people want from AI. It found that preferences vary widely, challenging current methods like RLHF that try to align AI with human values.

Researchers introduced a new benchmark, UnpredictaBench, to evaluate whether large language models (LLMs) can capture true underlying distributions rather than collapsing to a single plausible answer. This is critical as AI is increasingly used as a substitute for real entities in economic simulations and other modeling tasks.