
New Framework Predicts LLM Latency on Edge Devices with 95% Accuracy
Researchers have developed a runtime-aware framework that predicts how long large language models (LLMs) will take to respond on heterogeneous edge devices like smartphones and IoT gadgets, achieving 95% accuracy by accounting for hardware, runtime backend, thermal variation, and prompt behavior.


