← Search

Xiaokai Zhou

2 accepted papers

2026

Scheduling LLM Inference with Uncertainty-Aware Output Length Predictions

ICML 2026poster

To schedule LLM inference, the \textit{shortest job first} (SJF) principle is favorable by prioritizing requests with short output lengths to avoid head-of-line (HOL) blocking. Existing methods usually predict a single output length for each request to facilitate scheduling. We argue that such a \te…

Cited by 0SourceScholar
2025

Model Rake: A Defense Against Stealing Attacks in Split Learning

IJCAI 2025

Split learning is a prominent framework for vertical federated learning, where multiple clients collaborate with a central server for model training by exchanging intermediate embeddings. Recently, it is shown that an adversarial server can exploit the intermediate embeddings to train surrogate mode

Cited by 0SourcePDFScholar