← Search

Hongchao Zhu

1 accepted papers

2026

Scheduling LLM Inference with Uncertainty-Aware Output Length Predictions

ICML 2026poster

To schedule LLM inference, the \textit{shortest job first} (SJF) principle is favorable by prioritizing requests with short output lengths to avoid head-of-line (HOL) blocking. Existing methods usually predict a single output length for each request to facilitate scheduling. We argue that such a \te…

Cited by 0SourceScholar