OpenSkill: AI Agents That Learn Without Human Guidance
Researchers developed OpenSkill, a system that lets AI agents learn and improve on their own in the real world. This could make AI tools more adaptable and useful without constant human input.
1193 stories curated by AInformed · page 25 of 50
Researchers developed OpenSkill, a system that lets AI agents learn and improve on their own in the real world. This could make AI tools more adaptable and useful without constant human input.
OpenAI has introduced the Economic Research Exchange to explore how AI affects jobs, productivity, and the economy. Researchers can now apply to participate in selected projects.
A study found that AI personalization works poorly with real people. Researchers say current systems are often evaluated on fake data, not actual conversations. This could mean your AI assistant isn't really learning about you as well as it should.
Researchers developed a new method to detect AI hallucinations by analyzing relationships between evidence and answers. This approach could make AI responses more reliable for everyday users.
Scientists have discovered a new way to fine-tune AI models by separating the direction and strength of their internal signals. This could lead to more precise and reliable AI behavior. The research challenges the common assumption that only the direction of these signals matters, not their intensity.
Researchers developed a system that helps AI models save effort by quickly deciding whether a problem needs deep thought. This could make AI tools faster and more efficient for everyday users.
Researchers have developed a new method to reduce bias in AI systems by treating fairness as a symmetry operation. Tested on four synthetic datasets, the framework achieved upwards of 90% violation reduction, significantly improving fairness in high-stakes socioeconomic decisions like hiring or lending.
Researchers created a multilingual dataset to improve AI's ability to share facts across languages. This could make AI assistants more reliable worldwide.
Researchers developed a new memory system for AI that helps it learn from failures and improve over time. This could make AI assistants much better at complex, multi-step tasks.
Researchers introduced DiBS, a new AI method that merges diffusion models with traditional logic to solve Sudoku more efficiently. This hybrid approach aims to overcome the limitations of pure learning-based or symbolic solvers.
Researchers introduced Lean4Agent, a new method to make AI agents more reliable by using formal verification. This approach helps ensure AI agents follow correct steps in complex tasks, reducing errors in multi-step workflows.
Researchers created CrowdMath, a dataset of 164 expert-annotated math discussions showing how experts collaborate to solve open problems. This could help AI models learn from human reasoning processes, not just final answers.
Researchers introduced AEGIS, a system that uses a lightweight probe on a robot's activations to detect high-risk steps and switch to a stronger policy only when needed, preventing gradual failure in long-horizon manipulation tasks.
Researchers introduced TimeClaw, an agentic framework that improves time series analysis by integrating rich contextual information and supporting end-to-end workflows. This could make tools for forecasting and pattern analysis more accurate and practical for real-world applications.
Researchers have proposed a new motivational architecture for AI that focuses on conversational agents. The architecture reinterprets the OpenPsi motivational lineage for linguistic interactions, aiming to make future AI assistants more engaging and responsive to human mental states.
Researchers developed an AI system that predicts knee pain from MRI scans with high accuracy. The framework combines deep learning and statistical modeling to make the results interpretable and trustworthy.
Researchers developed a method called GITCO to improve AI forecasting by cleaning up the input data instead of changing the model itself. This could make AI predictions more accurate without needing to retrain the model.
When AI models edit code repeatedly, they tend to recycle the same solutions rather than exploring new ones. This could limit how creative AI tools can be when helping programmers. Researchers found that in 87% of mutation chains, over 93% of AI-generated code mutations revisited familiar structural forms.
Researchers analyzed a discontinued Reddit experiment where AI-powered accounts debated users without disclosure. The study highlights how these AI agents influenced opinions in a real-world setting.
Researchers introduced SentinelBench, a benchmark to test AI agents' ability to monitor tasks over long periods. This could improve AI assistants that handle slow, real-world tasks like waiting for stock price changes or tracking delivery updates.
Scientists have developed a new way to study how AI models can degrade when trained on synthetic data. Their findings show that this problem spreads between models, much like a contagious disease. This could help prevent future AI systems from becoming unreliable.
Scientists studied how AI agents communicate and found that unstructured chatting wastes resources. They discovered that structured communication can make AI teams work faster and cheaper.
Researchers introduced Agents' Last Exam (ALE), a new benchmark to evaluate AI agents on long-horizon, economically valuable tasks with verifiable outcomes. This could help bridge the gap between AI performance in labs and real-world usefulness.
Researchers created a synthetic dataset to help AI understand complex questions across multiple tables. This could make databases and spreadsheets much easier to query with natural language.