← Search

Yuanfan Chen

1 accepted papers

2026

Beyond Prediction: Tail-Aware Scheduling for LLM Inference

ICML 2026poster

LLM serving exhibits extreme length variability, making size-based scheduling difficult in practice. Recent LLM schedulers approximate SJF/SRPT using predicted decode lengths or rank and primarily report mean-centric metrics (e.g., TTFT/TBT). We show these prediction-driven policies can be fragile u…

Cited by 0SourceScholar