2026
MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Models Serving
ICML 2026poster
The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output le…