
Why Bigger AI Models Solve Harder Problems Better
Larger AI models consistently outperform smaller ones in reasoning tasks. Researchers developed a new tool to study why this happens, revealing key differences in problem-solving approaches.
936 stories tagged Research · page 11 of 39

Larger AI models consistently outperform smaller ones in reasoning tasks. Researchers developed a new tool to study why this happens, revealing key differences in problem-solving approaches.

A new research paper reveals that the classical intuition that verifying a solution is easier than producing one is being inverted for today's coding agents. As foundation models get stronger, generating candidate solutions has become easier, while reliable verification—capturing underspecified human intent—has become the harder problem.
Researchers argue that after a benchmark's accuracy saturates, the focus should shift to six other key dimensions of AI performance: construct validity (shortcuts), out-of-distribution generalizability, efficiency, reliability, model vs. scaffold importance, and human–AI collaboration uplift.

A new paper highlights flaws in how we test AI models that handle text, images, and other inputs together. Current methods miss key aspects like understanding physical reality or combining different types of information. The authors propose better evaluation frameworks to address these gaps.

Scientists found that editing one part of an AI's instructions can unintentionally change other parts. This happens because AI models share a common context window, causing unexpected behavior. This is a problem for developers who use AI to build complex systems.

Scientists discovered that AI chatbots refuse requests less often when they adopt a more cooperative personality. This finding could help make AI assistants more helpful while maintaining safety.

Scientists developed a method to identify and manage AI's tendency to flatter users. This breakthrough could make AI models more honest and reliable in everyday interactions.

Researchers developed a method to speed up advanced AI language models without needing to retrain them. This breakthrough could make AI text generation faster and more efficient for everyday use.

Researchers suggest a new way to govern AI agents by focusing on actions rather than their reasoning. This approach could make AI decisions more transparent and trustworthy in critical areas like healthcare and software deployment.

Researchers created OpenFinGym, a platform to test AI trading bots on multiple financial tasks. It helps evaluate how well these bots perform in real-world market conditions.

Researchers have developed HierBias, an AI model that analyzes whole articles to detect media bias, formally proving that using document context reduces error compared to sentence-by-sentence approaches.

Researchers developed an AI-powered pipeline to compare governance structures of decentralized and corporate AI protocols. The system analyzes over 4,300 governance discussions to uncover power dynamics in AI agent interoperability standards.

Researchers created a contamination-aware, multi-zone benchmark called Know2Guess to evaluate when large language models should answer questions versus abstain. It has 1,200 items across five domains with explicit abstention expectations and contamination-risk metadata.

Researchers found that AI models like GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 have distinct tendencies when suggesting research methods. The study highlights how these models might influence scientific research differently.

Researchers introduced COrigami, an AI pipeline that co-designs origami crease patterns which are both mathematically flat-foldable and visually recognizable. This bridges the gap between rigid geometric constraints and subjective aesthetics, potentially making computational origami design more accessible.

Researchers introduced ContextForge, a system that helps large language models maintain relevant information across long conversations by recycling context instead of replaying entire histories. This could slash token usage and improve multiturn reasoning.

Researchers developed an AI system that combines official drug data with patient experiences to provide safer, more accurate medication information. This could help people make better decisions about their mental health treatments.

Researchers have developed AlgoEvolve, an AI system that uses large language models to evolve and improve algorithmic trading strategies. This could make trading more accessible and efficient for everyday investors.

Researchers evaluated AI agents on complex energy market tasks, including live data retrieval, regulatory knowledge, and multi-step quantitative reasoning. The study fills a critical gap, with 243 expert-designed tasks showing how tool-augmented LLMs could make energy systems more efficient and responsive.

Researchers introduced TrustMem, a framework designed to improve long-term memory in AI agents. It addresses the problem of AI systems forgetting important information or generating hallucinated content that becomes permanently stored. This could make AI assistants more reliable over extended interactions.

A new AI model called iLLaDA uses a different approach to understand language, potentially making it more efficient and accurate. This could lead to better AI assistants and tools that understand context more naturally.

AI agents often make mistakes that snowball over time, especially in tasks like persuasion. Researchers found a key reason—semantic leakage in standard RAG—and developed a method called Taxonomic Strategy Retrieval to prevent these compounding errors.

A new practitioner's reference, 'The Hitchhiker's Guide to Agentic AI,' covers the full stack of building autonomous AI systems—from transformer architecture and GPU systems to fine-tuning, model compression, and production deployment.

Researchers developed an AI technique to automatically create highlights for academic papers. This could make scientific research easier to digest for non-experts and improve literature searches.