What Building Shippy Taught Us About Building AI Agents
AllenAI's open-source Shippy agent introduces modular components that can be swapped independently, making AI development more accessible and scalable for developers.
232 stories tagged AI Agents · page 3 of 10
AllenAI's open-source Shippy agent introduces modular components that can be swapped independently, making AI development more accessible and scalable for developers.
A new ArXiv study shows that LLM-based AI agents lose information when communicating via plain text, suggesting they possess an internal 'world model' that exceeds textual expressibility. This finding has implications for multi-agent system design and alignment.
A study of 157 enterprises reveals that half have shipped AI agents that passed internal evaluations but failed in production. Only 5% fully trust automated evaluations, yet two-thirds are moving toward fully automated deployments without human oversight. The core issue is not test coverage but a reality-alignment gap between evaluations and real-world outcomes.
A VentureBeat Pulse Research report of 101 enterprises finds that most deployed 'AI agents' are actually simple chatbot wrappers, not true autonomous agents. Anthropic's Claude leads agent orchestration adoption due to reliable multi-step execution, while enterprises deliberately build hybrid control planes to avoid vendor lock-in. Real-time fiscal control over token burn remains rare.
A VentureBeat survey of 107 enterprises reveals that 54% have already experienced an AI agent security incident or near-miss. Despite this, most companies still let agents share credentials, only a third give each agent its own scoped identity, and just 30% isolate their highest-risk agents. The security stack is largely borrowed from model providers and hyperscalers rather than purpose-built for agents, and spending remains a thin slice of the security budget.
A new study reveals that the format of messages between AI agents significantly impacts how accurately information is preserved across multiple hops. Structured messages help maintain fidelity, but the effect depends on the type of information and the agent's tier.
Researchers created a new test called Long-Horizon-Terminal-Bench to evaluate AI agents on complex, time-consuming tasks. Unlike previous tests, it measures progress over time, not just final results.
Researchers developed a new AI planning method called GATS that reduces computational costs and improves reliability by eliminating LLM calls during inference. It could make AI assistants more efficient at complex tasks.
NVIDIA has released a massive open-source dataset designed to train AI agents. This data could accelerate the development of more capable AI assistants for everyday tasks.
AI agents are transforming banking at record speed, outpacing even the internet's adoption. The market for AI in banking is projected to reach $97B by 2027 and over $200B beyond, with the token $SERV positioned at the intersection of AI agents and banking.
Researchers discovered a new technique called 'Ghostcommit' that embeds hidden instructions in images to manipulate AI systems. This could trick AI agents into revealing sensitive information or performing unauthorized actions.
Lyzr, a startup building AI agents for businesses, used its own AI agent to successfully raise $100 million. This demonstrates the agent's capabilities in real-world applications.
Researchers found that AI agents can create reusable workflows from basic tasks, reducing errors and saving time. This could make AI assistants much more efficient at complex jobs.
Researchers introduced AgentLens, a benchmark that evaluates AI coding assistants by analyzing their entire problem-solving process. This goes beyond just checking if the code works, looking at how the AI follows instructions and recovers from mistakes.
A new study identifies recurring weaknesses in AI agents that use tools and plan tasks. These failures highlight challenges in making AI more reliable for everyday use.
Google has added new features to its Gemini API, including background tasks and remote MCP, making it easier for developers to build reliable AI agents. This could lead to more sophisticated AI assistants for everyday users.
A new paper explores integrating memory into every step of an AI agent's reasoning loop, but warns this approach could inflate latency by up to 83x. The work highlights a tension between memory-rich reasoning and real-time performance.
Mkrrm is a new AI tool that doesn't just research but actually completes tasks by using its own browser. It runs locally on your machine, keeping your data private and secure.
A security flaw in Cursor's AI agent allowed it to break out of its protective sandbox. This shows why AI tools need strict boundaries to prevent misuse. Developers are now working on fixes to ensure safer AI interactions.
Vercel's CEO argues that separating AI models from agents is crucial for better performance and cost efficiency. This shift could make AI tools faster and more affordable for businesses and developers.
A new open-source project called AgentLine allows AI agents to make phone calls. The project provides the infrastructure for AI systems to handle voice interactions, call management, and telephony integration, potentially enabling AI assistants to handle real-world tasks like scheduling appointments or customer service calls.
Speck is a new open-source tool that lets you define AI agents with simple text files, making it easier to automate complex tasks. It's designed to be as reliable as traditional software tools, but with the flexibility of AI.
Mark Zuckerberg told Meta employees that AI development isn't advancing as fast as hoped. This could delay consumer-facing AI products and services.
Researchers propose a new method to prevent AI agents from acting against user intentions. This approach could make AI tools safer and more trustworthy by tracking their actions like a digital paper trail.