2026
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
ICML 2026poster
Serving Large Language Models (LLMs) can benefit immensely from parallelizing both the model and input requests across multiple devices, but incoming workloads exhibit substantial spatial and temporal heterogeneity. Spatially, workloads comprise heterogeneous requests with varying compute and memory…