
Why Bigger AI Models Solve Harder Problems Better
Larger AI models consistently outperform smaller ones in reasoning tasks. Researchers developed a new tool to study why this happens, revealing key differences in problem-solving approaches.
1496 stories tagged AI · page 20 of 63

Larger AI models consistently outperform smaller ones in reasoning tasks. Researchers developed a new tool to study why this happens, revealing key differences in problem-solving approaches.

A new research paper reveals that the classical intuition that verifying a solution is easier than producing one is being inverted for today's coding agents. As foundation models get stronger, generating candidate solutions has become easier, while reliable verification—capturing underspecified human intent—has become the harder problem.
Researchers argue that after a benchmark's accuracy saturates, the focus should shift to six other key dimensions of AI performance: construct validity (shortcuts), out-of-distribution generalizability, efficiency, reliability, model vs. scaffold importance, and human–AI collaboration uplift.

The Trump administration has asked OpenAI to release its advanced AI models in stages rather than all at once, according to a new report. This request reflects growing government concern about the rapid pace of AI development and its potential risks.

Scientists discovered that AI chatbots refuse requests less often when they adopt a more cooperative personality. This finding could help make AI assistants more helpful while maintaining safety.

OpenAI has unveiled GPT-5.6 Sol, a new AI model that excels in coding, science, and cybersecurity. It also includes advanced safety features to prevent misuse. This could revolutionize how professionals use AI in technical fields.

Researchers developed a method to speed up advanced AI language models without needing to retrain them. This breakthrough could make AI text generation faster and more efficient for everyday use.

Researchers created OpenFinGym, a platform to test AI trading bots on multiple financial tasks. It helps evaluate how well these bots perform in real-world market conditions.

Researchers have developed HierBias, an AI model that analyzes whole articles to detect media bias, formally proving that using document context reduces error compared to sentence-by-sentence approaches.

Researchers created a contamination-aware, multi-zone benchmark called Know2Guess to evaluate when large language models should answer questions versus abstain. It has 1,200 items across five domains with explicit abstention expectations and contamination-risk metadata.

Researchers found that AI models like GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 have distinct tendencies when suggesting research methods. The study highlights how these models might influence scientific research differently.

Researchers introduced COrigami, an AI pipeline that co-designs origami crease patterns which are both mathematically flat-foldable and visually recognizable. This bridges the gap between rigid geometric constraints and subjective aesthetics, potentially making computational origami design more accessible.

Researchers introduced ContextForge, a system that helps large language models maintain relevant information across long conversations by recycling context instead of replaying entire histories. This could slash token usage and improve multiturn reasoning.

Researchers developed an AI system that combines official drug data with patient experiences to provide safer, more accurate medication information. This could help people make better decisions about their mental health treatments.

Researchers have developed AlgoEvolve, an AI system that uses large language models to evolve and improve algorithmic trading strategies. This could make trading more accessible and efficient for everyday investors.

A new AI model called iLLaDA uses a different approach to understand language, potentially making it more efficient and accurate. This could lead to better AI assistants and tools that understand context more naturally.

OpenAI and Broadcom have developed a new chip designed to make AI models like ChatGPT run faster and cost less. This could lead to more powerful and affordable AI tools for everyone.

OpenAI and Broadcom have introduced Jalapeño, a custom AI chip built specifically for LLM inference. The chip aims to improve performance, efficiency, and scalability of AI systems, potentially making AI tools faster and more cost-effective for everyday users.

NVIDIA's new NeMo AutoModel makes customizing AI models faster and easier. It automates the fine-tuning process, which could help businesses and developers create specialized AI tools without needing deep technical expertise.

CtxGov is a new open-source tool that reveals the hidden instructions behind AI agents. It helps users understand how their AI assistants make decisions. This transparency could change how people trust and use AI tools.

A new practitioner's reference, 'The Hitchhiker's Guide to Agentic AI,' covers the full stack of building autonomous AI systems—from transformer architecture and GPU systems to fine-tuning, model compression, and production deployment.

Researchers developed an AI technique to automatically create highlights for academic papers. This could make scientific research easier to digest for non-experts and improve literature searches.

A small group of Wikipedia editors has shown that targeted edits can influence how AI models discuss animal welfare. Their work highlights the power of curated information in shaping AI behavior.

Ford is bringing back 350 quality inspectors after AI tools failed to capture their expertise or train new employees effectively. This highlights the challenges of replacing human knowledge with automation in complex fields.