AI Agents Learn to Build Their Own Reusable Workflows
Researchers found that AI agents can create reusable workflows from basic tasks, reducing errors and saving time. This could make AI assistants much more efficient at complex jobs.
1100 stories tagged Research · page 14 of 46
Researchers found that AI agents can create reusable workflows from basic tasks, reducing errors and saving time. This could make AI assistants much more efficient at complex jobs.
Researchers introduced AgentLens, a benchmark that evaluates AI coding assistants by analyzing their entire problem-solving process. This goes beyond just checking if the code works, looking at how the AI follows instructions and recovers from mistakes.
A new study identifies recurring weaknesses in AI agents that use tools and plan tasks. These failures highlight challenges in making AI more reliable for everyday use.
Researchers created a new system to train AI agents in realistic simulations, using reinforcement learning and reward shaping to improve multi-step decision-making.
Researchers have developed an AI system that helps non-experts design complex industrial parts using simple language. This tool could make professional-level design more accessible to everyone.
Researchers have developed an AI system called Prompt-to-Paper that creates scientific papers from prompts, ensuring claims are grounded in real literature. This addresses issues like fabricated results and lack of quality standards in AI-generated research.
Researchers have developed a new method to uncover the underlying causes of AI decisions in cyber-physical systems. This approach provides more robust insights, helping users understand automated decisions, especially in high-risk domains.
Researchers developed a new AI model called Narrative World Model (NWM) to help writers manage complex story details. It keeps track of evolving story states, like character secrets and event timelines, to improve long-form fiction writing.
Researchers introduced FirstResearch, a framework that generates a structured Research Question Certificate for AI-suggested scientific questions. The certificate records primitive definitions, assumptions, and falsifiers, allowing scientists to audit the reasoning behind each question.
Researchers introduced Akashic, a low-overhead memory system for LLM inference that uses MemAttention to organize context into bounded chunks and model semantic relationships. This could make chatbots and AI assistants much more efficient and accurate by reducing prefill costs and avoiding context limits.
A new paper explores integrating memory into every step of an AI agent's reasoning loop, but warns this approach could inflate latency by up to 83x. The work highlights a tension between memory-rich reasoning and real-time performance.
Researchers developed VERITAS, an AI framework to automatically replicate scientific studies, addressing the growing challenge of verifying published research. This could make it easier to check the accuracy of scientific findings, benefiting both researchers and the public.
Researchers developed Oyster-II, an AI safety system that helps models answer sensitive questions more constructively. It improves on previous approaches by providing useful information instead of just refusing requests.
Researchers developed REDI, a framework to automate the complex process of preparing scientific data for AI training. It handles everything from data cleanup to tracking where the data came from, making it easier for scientists to use AI effectively.
Researchers introduced a new approach to make AI systems better understand and align with human decision-making. This could help people trust and use AI tools more effectively in everyday tasks.
Neuronpedia is an open-source platform that makes AI models more transparent by visualizing their inner workings. It helps researchers and developers understand how AI systems make decisions.
Researchers have developed a new unified multimodal foundation model that jointly models vision, language, world dynamics, and action generation. This advancement could lead to more capable robots and virtual assistants that can understand instructions, anticipate environmental changes, and execute precise actions over extended horizons.
Researchers have unveiled Gemma 4, a new suite of open-weight AI models that handle text, images, and audio together. These models are designed to be more efficient and better at reasoning, with sizes ranging from 2.3 billion to 31 billion parameters.
Researchers developed a new AI approach that mimics how doctors gather evidence to make diagnoses. This method could make AI medical tools more accurate and reliable.
AI systems are vulnerable to 'prompt injection' attacks, where hackers trick them into revealing data or performing unauthorized actions. A new survey examines the best ways to protect against these threats.
Researchers at Thinking Machines AI have published a study demonstrating an AI system that can replicate expert financial decision-making. The system was trained on data from seasoned financial analysts and shows proficiency in market trend analysis, risk assessment, and strategic planning. This could make sophisticated financial advice more accessible to everyday investors.
Google's NotebookLM is adding a new feature that creates 60-second AI videos summarizing your research. This is rolling out to Google AI Ultra and Pro subscribers, making it easier to digest complex information quickly.
Researchers have introduced Wiola, a novel small language model architecture that breaks from existing designs. It incorporates five unique components aimed at improving efficiency and performance.
Researchers developed TokenScope, a tool that reveals how AI models make decisions when writing code. It provides real-time insights into the AI's thought process, helping developers understand and trust AI-generated code.