← Search

Wenhai Lin

1 accepted papers

2026

AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving

ICML 2026poster

Augmented large language models (LLMs) that invoke external calls are increasingly prevalent in inference serving. However, such augmentations pose significant challenges to inference efficiency under strict Service-Level Objectives (SLOs). Existing inference systems are agnostic to the dynamic exec…

Cited by 0SourceScholar