Why AI Agents Could Be the Key to Scalable Enterprise Adoption
IBM Research explains why AI agents—software that automates complex tasks—are crucial for businesses to use AI at scale. These agents could make AI more practical for everyday work.
255 stories tagged LLMs · page 4 of 11
IBM Research explains why AI agents—software that automates complex tasks—are crucial for businesses to use AI at scale. These agents could make AI more practical for everyday work.
Scientists developed a new AI model called CSRM that can be quickly adjusted to meet changing safety requirements. This could help AI systems better adapt to new rules and regulations as they evolve.
Researchers introduced PReMISE, a framework that uses pairwise human-preference data to discover policy-level rubric sets and audit LLM judges, ensuring AI scores align with human preferences and avoiding misleadingly polished but factually incorrect responses.
Researchers developed a method called AdaCoM to help AI agents manage information overload during long tasks. Unlike prior approaches that require retraining the agent itself, AdaCoM adapts context management strategies to each agent, making it practical for closed-source systems.
Researchers introduced DecomposeR, a new AI framework that improves how large language models plan and execute complex research tasks. By representing research plans as typed directed acyclic graphs (DAGs), DecomposeR enables better credit assignment for planning and execution, potentially leading to more structured and accurate long-form answers.
Researchers propose a new way for AI to automatically gather and prepare its own specialized training data. This could make AI models better at niche tasks without human help. The approach, called 'Autonomous Agentic Data Engineering,' lets AI act as its own data engineer, potentially speeding up customization for fields like medicine or law.
A tech hobbyist successfully ran a massive AI language model using 768GB of Intel Optane memory and a single graphics card. The setup, called Local Kimi K2.5, achieved roughly 4 tokens per second, showing how powerful consumer hardware can be for cutting-edge AI experiments.
Tiny-vLLM is a new open-source tool that makes running large language models faster and more efficient. It could make advanced AI tools more accessible to regular users.
Scientists have created StoryMI, a multi-AI system that can generate realistic therapeutic dialogues. This tool could help train therapists and provide new ways to practice motivational interviewing techniques.
Researchers found that different large language models (LLMs) often categorize the same public comments in conflicting ways. This can shape what policymakers see, potentially biasing decisions. The study proposes a new audit pipeline to flag disagreements for human review.
Kog AI has developed a method to run large language models at 3,000 tokens per second on standard GPUs, making advanced AI faster and more accessible. This breakthrough could significantly reduce the cost and complexity of AI applications.
A detailed breakdown explains how large language models like ChatGPT are built and function. This helps demystify the technology for curious beginners.
Researchers have developed a new AI system that can detect human values in text, both explicit and implicit. This could help make AI decisions more aligned with ethical and moral considerations.
A new study highlights how AI models can accidentally expose sensitive training data. This raises privacy concerns and challenges for AI developers. Researchers propose ways to mitigate these risks.
Researchers created a benchmark to test AI's ability to model human beliefs and emotions. This could help build more empathetic and socially aware AI assistants.
A new study challenges claims that AI models can introspect, arguing that what looks like self-awareness might just be clever pattern matching. Researchers say we need better tests to truly know if AI understands itself.
A new study shows that AI models with strong medical knowledge can abandon correct diagnoses when pressured in conversations. Researchers developed a test to measure how well AI maintains accurate beliefs under stress.
Norway is deploying 2 petabytes of Huawei flash storage to support large language model (LLM) training. This move highlights the growing demand for high-speed storage solutions in AI development.
A new study categorizes AI sycophancy into clear types, helping developers build more honest chatbots. The research highlights how current AI models often agree with users even when they're wrong, making conversations less reliable.
Researchers created a benchmark called MOOD to test AI models' ability to detect unexpected safety failures. This could help prevent AI from behaving dangerously in unusual situations.
A new guide explains essential data concepts in simple terms, helping beginners grasp how AI models like LLMs work. This primer makes complex ideas accessible without requiring a technical background.
Researchers developed a new method called OSCToM to help AI models better understand nested beliefs and conflicts in social settings. This could make AI assistants more adept at navigating complex human interactions.
Researchers found that AI models often fail to handle rare medical cases not covered by standard guidelines. This highlights a critical gap in how medical AI is trained and evaluated.
Scientists suggest developing data probes to better understand how different types of information affect AI models. This could make training large language models more efficient and effective.