New Research Reveals How to Detect and Control AI's Sycophantic Behavior
Scientists developed a method to identify and manage AI's tendency to flatter users. This breakthrough could make AI models more honest and reliable in everyday interactions.
155 stories tagged Machine Learning · page 3 of 7
Scientists developed a method to identify and manage AI's tendency to flatter users. This breakthrough could make AI models more honest and reliable in everyday interactions.
A new AI model called iLLaDA uses a different approach to understand language, potentially making it more efficient and accurate. This could lead to better AI assistants and tools that understand context more naturally.
Researchers developed Heuresis, a framework to help AI agents explore high-quality, diverse, and novel ideas in machine learning. This could speed up scientific progress by making AI research more efficient and creative.
General Intuition raised $320 million in fresh funding—bringing its total raised to $2.3 billion—to train AI models on millions of hours of video game gameplay, aiming to teach AI something closer to human intuition.
Researchers developed PEAR, a dynamic AI debate system that improves reliability by changing roles and reducing biases. This could make AI responses more trustworthy and consistent.
A new research paper introduces a reference architecture for 'agent skills' — reusable, externalized behavioral knowledge that LLM agents can discover, activate, and interpret at runtime. The framework formalizes how skills are bound to context and authority, interpreted by stochastic agents, and recorded as run evidence.
A new paper proposes a roadmap for AI that learns through interaction with complex environments, mimicking natural evolution by systematically removing human priors. The authors have released an open-source infrastructure called Darwin Mobile Agent to test this idea using mobile graphical user interfaces as a practical proxy for an open-ended world.
Researchers have developed AlphaMemo, an AI system that helps financial trading agents learn from their past decisions. This could make AI-driven trading more reliable and less prone to repeating errors.
Scientists developed a new method for AI to better distinguish between different types of uncertainty — whether it's missing information or just inherent randomness. This could make LLMs more reliable in real-world tasks where mistakes matter.
Researchers found that selecting the most informative comparison pairs during AI training can improve results. This approach could make AI models more efficient to train without extra cost.
Researchers at Rice University have developed a system that allows robots to learn complex tasks by watching videos. This could make robots more adaptable and useful in everyday environments.
Researchers have developed new open-source methods that outperform LoRA, the most popular AI fine-tuning technique. These approaches could make customizing large language models more accessible and affordable for everyone.
Researchers used AI coding agents to autonomously direct robots through complex tasks like installing GPUs and cutting zip-ties. This breakthrough could reduce the need for manual programming in manufacturing and other industries.
Researchers developed a new AI framework called SERAF that enhances time series forecasting by combining historical patterns with semantic context to address non-stationarity. This approach could make predictions more accurate for real-world applications like stock markets and weather forecasting.
Researchers developed a new type of AI model that can reason about unseen combinations of objects by combining causal and relational reasoning. This breakthrough could help AI systems generalize better to new situations.
A new study demonstrates that AI models can pick up subtle biases or 'vibes' from training data without being directly taught. This raises important questions about fairness in AI systems.
PyTorch has introduced a new optimization technique called fused MLP that speeds up AI model training. This advancement makes it easier and faster for developers to build and train complex AI models.
Hugging Face now offers a new way to run continuous integration (CI) workflows directly on its platform. This can simplify the process for developers using GitHub repositories.
New research shows AI models' self-reported traits often don't match their actual behavior, but this may be due to weak testing methods—not incoherent AI. The study calls for more specific, context-matched psychometric probes to better predict AI actions in real-world deployments.
Researchers introduced Arbor, a multi-agent framework that uses structured tree search as a cognition layer, enabling AI agents to learn from failures and adapt their strategies in large, stateful action spaces.
Researchers developed CodeAlchemy, a system that generates synthetic code to train AI models. This could make AI coding tools smarter and more versatile for real-world tasks.
Researchers have developed a new method to reduce bias in AI systems by treating fairness as a symmetry operation. Tested on four synthetic datasets, the framework achieved upwards of 90% violation reduction, significantly improving fairness in high-stakes socioeconomic decisions like hiring or lending.
Researchers developed a method called GITCO to improve AI forecasting by cleaning up the input data instead of changing the model itself. This could make AI predictions more accurate without needing to retrain the model.
Researchers introduced Grokers, an AI system that builds structured understanding of knowledge graphs by analyzing data upfront. This could make AI interactions faster and more accurate for complex queries.