2025
LLM Query Scheduling with Prefix Reuse and Latency Constraints
NeurIPS 2025poster
The efficient deployment of large language models (LLMs) in online settings requires optimizing inference performance under stringent latency constraints, particularly the time-to-first-token (TTFT) and time-per-output-token (TPOT). This paper focuses on the query scheduling problem for LLM inferenc…