
New AI Memory Framework Helps LLMs Learn from Mistakes
Researchers developed a new memory system for AI that helps it learn from failures and improve over time. This could make AI assistants much better at complex, multi-step tasks.
238 stories tagged LLMs · page 3 of 10

Researchers developed a new memory system for AI that helps it learn from failures and improve over time. This could make AI assistants much better at complex, multi-step tasks.

Researchers analyzed a discontinued Reddit experiment where AI-powered accounts debated users without disclosure. The study highlights how these AI agents influenced opinions in a real-world setting.

Researchers introduced SMAC-Talk, a new AI challenge that tests how well LLM-based agents communicate and coordinate in complex, partially-observable environments. This could help build AI systems that work together effectively in real-world scenarios like disaster response or smart cities.

A new study proposes Visual Graph Scaffolds, a method that uses graph structures to improve the reasoning of large language models (LLMs). Unlike prior approaches that treat graphs as external data sources, this technique integrates graphs directly into the model's reasoning process, inspired by how humans use mind maps to organize complex thoughts. The method showed significant improvements on multi-hop question answering tasks.

Researchers created ChatHealthAI to combine AI's language skills with medical records. This could help doctors make better decisions using patient history. The system is still in early development, but it shows promise for improving healthcare.

Researchers created a multi-domain red teaming framework to evaluate how well AI models handle complex medical scenarios. The study found significant gaps in safety and fairness across 11 leading AI models.

Researchers created AbaqusAgent, a multi-agent AI system that automates finite element analysis (FEA) for engineering. This could make advanced simulations accessible to non-experts, speeding up product design and reducing errors.

IBM Research explains why AI agents—software that automates complex tasks—are crucial for businesses to use AI at scale. These agents could make AI more practical for everyday work.

Scientists developed a new AI model called CSRM that can be quickly adjusted to meet changing safety requirements. This could help AI systems better adapt to new rules and regulations as they evolve.

Researchers introduced PReMISE, a framework that uses pairwise human-preference data to discover policy-level rubric sets and audit LLM judges, ensuring AI scores align with human preferences and avoiding misleadingly polished but factually incorrect responses.

Researchers developed a method called AdaCoM to help AI agents manage information overload during long tasks. Unlike prior approaches that require retraining the agent itself, AdaCoM adapts context management strategies to each agent, making it practical for closed-source systems.

Researchers introduced DecomposeR, a new AI framework that improves how large language models plan and execute complex research tasks. By representing research plans as typed directed acyclic graphs (DAGs), DecomposeR enables better credit assignment for planning and execution, potentially leading to more structured and accurate long-form answers.

Researchers propose a new way for AI to automatically gather and prepare its own specialized training data. This could make AI models better at niche tasks without human help. The approach, called 'Autonomous Agentic Data Engineering,' lets AI act as its own data engineer, potentially speeding up customization for fields like medicine or law.

A tech hobbyist successfully ran a massive AI language model using 768GB of Intel Optane memory and a single graphics card. The setup, called Local Kimi K2.5, achieved roughly 4 tokens per second, showing how powerful consumer hardware can be for cutting-edge AI experiments.

Tiny-vLLM is a new open-source tool that makes running large language models faster and more efficient. It could make advanced AI tools more accessible to regular users.

Scientists have created StoryMI, a multi-AI system that can generate realistic therapeutic dialogues. This tool could help train therapists and provide new ways to practice motivational interviewing techniques.

Researchers found that different large language models (LLMs) often categorize the same public comments in conflicting ways. This can shape what policymakers see, potentially biasing decisions. The study proposes a new audit pipeline to flag disagreements for human review.

Kog AI has developed a method to run large language models at 3,000 tokens per second on standard GPUs, making advanced AI faster and more accessible. This breakthrough could significantly reduce the cost and complexity of AI applications.

A detailed breakdown explains how large language models like ChatGPT are built and function. This helps demystify the technology for curious beginners.

Researchers have developed a new AI system that can detect human values in text, both explicit and implicit. This could help make AI decisions more aligned with ethical and moral considerations.

A new study highlights how AI models can accidentally expose sensitive training data. This raises privacy concerns and challenges for AI developers. Researchers propose ways to mitigate these risks.

Researchers created a benchmark to test AI's ability to model human beliefs and emotions. This could help build more empathetic and socially aware AI assistants.

A new study challenges claims that AI models can introspect, arguing that what looks like self-awareness might just be clever pattern matching. Researchers say we need better tests to truly know if AI understands itself.

A new study shows that AI models with strong medical knowledge can abandon correct diagnoses when pressured in conversations. Researchers developed a test to measure how well AI maintains accurate beliefs under stress.