
Codex is becoming a productivity tool for everyone
OpenAI's Codex is evolving beyond coding to help with research, data analysis, and content creation. It's making complex tasks easier for everyday users, not just programmers.
936 stories tagged Research · page 19 of 39

OpenAI's Codex is evolving beyond coding to help with research, data analysis, and content creation. It's making complex tasks easier for everyday users, not just programmers.

When AI models edit code repeatedly, they tend to recycle the same solutions rather than exploring new ones. This could limit how creative AI tools can be when helping programmers. Researchers found that in 87% of mutation chains, over 93% of AI-generated code mutations revisited familiar structural forms.

Researchers introduced SentinelBench, a benchmark to test AI agents' ability to monitor tasks over long periods. This could improve AI assistants that handle slow, real-world tasks like waiting for stock price changes or tracking delivery updates.

Scientists have developed a new way to study how AI models can degrade when trained on synthetic data. Their findings show that this problem spreads between models, much like a contagious disease. This could help prevent future AI systems from becoming unreliable.

Scientists studied how AI agents communicate and found that unstructured chatting wastes resources. They discovered that structured communication can make AI teams work faster and cheaper.

Researchers introduced Agents' Last Exam (ALE), a new benchmark to evaluate AI agents on long-horizon, economically valuable tasks with verifiable outcomes. This could help bridge the gap between AI performance in labs and real-world usefulness.

Researchers created a synthetic dataset to help AI understand complex questions across multiple tables. This could make databases and spreadsheets much easier to query with natural language.

Researchers developed a new AI system called Query Retrieve Conclude that can understand and interpret memes by finding missing context online. This could help platforms moderate content and users better understand internet humor.

Researchers introduced LeanMarathon, a multi-agent AI system designed to help mathematicians formalize and prove complex theorems in the Lean proof assistant. It uses four contract-scoped agents to construct, audit, prove, and repair an evolving blueprint that serves as a formal proof skeleton, natural-language proof graph, and shared system of record, addressing issues like statement drift, tangled dependencies, and context decay.

EVA-Bench Data 2.0 is a comprehensive dataset designed to test AI models' ability to use tools effectively. It includes 213 scenarios across 3 domains and 121 tools, making it a valuable resource for developers and researchers.

Researchers discovered that AI judges used to rank model performance can be swayed by follow-up conversations after they have already made a decision. This vulnerability, called 'post-decision manipulability,' challenges the reliability of current AI evaluation methods.

Researchers introduced SMAC-Talk, a new AI challenge that tests how well LLM-based agents communicate and coordinate in complex, partially-observable environments. This could help build AI systems that work together effectively in real-world scenarios like disaster response or smart cities.

Researchers developed PEEL, a new framework to improve transparency in AI-assisted research. It combines text analysis tools with AI interpretation to spot distortions in research findings.

Researchers propose a way to test AI agents before they go live, ensuring they follow rules, stay safe, and comply with governance standards. This addresses a critical gap in enterprise AI reliability.

Researchers created VAMPS (Visual-Assisted Mathematical Problem Solving), a benchmark to test AI models' ability to solve math problems using visual tools like graphs. This is important because real-world science and engineering often rely on visual aids for problem-solving, and many current AI models struggle when they must use external tools and interpret their visual outputs.

Researchers developed StepPRM-RTL, an AI framework that enhances the accuracy of automatically generating hardware code. This could make designing digital circuits faster and more reliable for engineers.

A new paper argues emotional support from AI often arises incidentally during routine, task-oriented interactions—not just from dedicated companion chatbots—and this incidental bonding could reshape how people connect with both machines and humans.

Researchers have extended Direct Preference Optimization (DPO) to improve AI models beyond just chatbots. This could make AI assistants more helpful and safer across various applications. Researchers open-sourced the code, making it accessible to developers and researchers worldwide.

Researchers created a benchmark to test if AI agents can handle the labor-intensive task of curating training data for other AI systems. This could drastically speed up AI development by automating a key bottleneck.

Researchers found that AI tools are changing how mathematicians formalize and verify proofs. These tools make it easier for humans to turn abstract ideas into machine-checked proofs, speeding up the process significantly.

Researchers argue that disagreements among AI systems might reveal important normative uncertainties, not just errors. They propose a knowledge-representation layer that abstracts reasoning traces and decisions into symbolic disagreement states for value-laden tasks.

AI systems often act without proper authorization or evidence, a problem called 'compliance bias'. Researchers propose new ways to evaluate when AI should abstain from actions. This could make AI safer and more reliable in real-world use.

A new study proposes Visual Graph Scaffolds, a method that uses graph structures to improve the reasoning of large language models (LLMs). Unlike prior approaches that treat graphs as external data sources, this technique integrates graphs directly into the model's reasoning process, inspired by how humans use mind maps to organize complex thoughts. The method showed significant improvements on multi-hop question answering tasks.

Researchers created BehaviorBench, a new AI benchmark that uses real-world behavioral data to test how well AI systems can personalize decisions. This could lead to more tailored AI assistants that understand individual preferences better.