MERIT Benchmark Reveals When Long-Term Memory Actually Helps AI Agents — and When It Doesn't
A new benchmark called MERIT measures the marginal utility of long-term memory for tool-using LLM agents. The study finds that memory often fails to improve task performance, especially when access costs are high, challenging assumptions about AI memory in real-world applications.