
TTKV: A New Approach to Optimizing Long-Context LLM Inference
Researchers propose TTKV, a temporal-tiered KV cache that prioritizes recent memories in LLMs, improving efficiency for long-context inference. This method mimics human memory systems, offering a more scalable solution than existing approaches.