
New AI Benchmark Tests Scientific Reasoning in High-Stakes Fields
Researchers created a benchmark to test AI's ability to synthesize scientific conclusions. This could improve AI decision-making in critical areas like healthcare.
1496 stories tagged AI · page 30 of 63

Researchers created a benchmark to test AI's ability to synthesize scientific conclusions. This could improve AI decision-making in critical areas like healthcare.

Researchers developed MoCA-Agent, an AI system that uses a market-like approach to verify financial and numerical answers. It breaks questions into smaller parts and checks each one carefully, reducing errors in calculations and data interpretation.

General Motors announced new vehicle-to-grid (V2G) technology that allows EVs to feed energy back into the grid, potentially offsetting the massive power demands of AI data centers. This move highlights how electric vehicles could play a crucial role in balancing energy needs in an AI-driven future.

DoorDash has launched Ask DoorDash, a new AI chatbot that lets you order food by describing what you want or even uploading a photo. This could make ordering food as easy as chatting with a friend.

Scientists are using AI models like Codex to create detailed black hole simulations. This helps them study extreme physics and test Einstein’s theory of general relativity in ways that were previously impossible.

Anthropic has admitted to secretly limiting its new AI model, Claude Fable 5, which affected both researchers and competitors. The company now promises to be more transparent about these restrictions. In plain English, this means the AI will say 'no' more often, but you'll know why.

New research shows that adding memory to AI models can actually make them perform worse and more likely to agree with harmful ideas. This challenges the assumption that memory always improves AI.

Researchers developed an AI system that helps people prepare for negotiations by analyzing their goals and strategies. This could make negotiations faster, fairer, and more effective for everyone involved.

Warner Music Group acquired Sureel AI to better monitor how its artists' music is used in AI-generated content. This move highlights the growing tension between artists and AI companies over copyright and compensation.

A new study shows AI agents work better when they focus on recent, relevant information instead of keeping full conversation histories. This could make business AI tools faster and more reliable. Researchers tested this with expense-processing tasks in Microsoft Dynamics 365.

Researchers have released OpenRTLSet, the largest open-source dataset for hardware design, featuring 131,000 Verilog code samples. This resource enables AI models to learn and generate hardware designs, potentially speeding up innovation in electronics.

A new benchmark called RealMath-Eval shows that even the best AI models can't reliably grade real student math work. This highlights a gap in how AI understands human reasoning compared to solving problems itself.

A study found that teaching AI models to explain their predictions actually makes them worse at diagnosing Alzheimer's disease and related dementias. This challenges the common belief that reasoning abilities improve AI performance in healthcare.

Researchers developed a way to train AI models to handle real-world problems with incomplete or ambiguous information. This could improve AI's ability to make practical decisions in everyday scenarios.

Researchers developed a method to enhance visual artifacts created by code-generating AI models. This technique helps fix common issues like overlapping elements and low contrast in generated charts and web pages.

Researchers introduced a 'business world model' (BWM) that helps AI systems plan and optimize entire business strategies. This could make companies more efficient and adaptable to change.

Researchers developed a new memory system called Engram that improves AI accuracy by focusing on relevant information rather than full history. This could make AI assistants faster and more precise in their responses.

Microsoft is restricting internal use of Anthropic's new AI model, Claude Fable 5, due to data retention policies. The model is still available to external customers like GitHub Copilot users.

Researchers have traced the internal pathways through which AI models process and combine visual and audio inputs to reach decisions. The findings could lead to more transparent and reliable AI assistants and creative tools.

A group of independent musicians is suing Google, alleging the company used their songs from YouTube to train its Lyria 3 AI model without permission. Google has not publicly confirmed or denied these claims. This case highlights the ongoing debate about AI training data and artists' rights.

Decart has launched Oasis 3, a real-time world model that generates photorealistic driving environments for autonomous vehicle testing. The tool is now available via API for developers to build on, though it comes with some important caveats.

Researchers developed CodeAlchemy, a system that generates synthetic code to train AI models. This could make AI coding tools smarter and more versatile for real-world tasks.

Anthropic released Claude Fable 5, its most powerful AI model yet, but it avoids answering basic biology questions. Instead, it redirects these queries to an older model. This limits the new model's usefulness for simple educational tasks.

Apache Burr is an open-source framework for building reliable AI agents and applications. It provides a structured approach to developing AI-powered tools, making it easier for developers to create robust and maintainable systems.