← Search

Minlan Yu

2 accepted papers

2025

DON’T STOP ME NOW: EMBEDDING BASED SCHEDULING FOR LLMS

ICLR 2025poster

Efficient scheduling is crucial for interactive Large Language Model (LLM) applications, where low request completion time directly impacts user engagement. Size-based scheduling algorithms like Shortest Remaining Process Time (SRPT) aim to reduce average request completion time by leveraging known…

Cited by 4SourcePDFScholar
2025

Fast Inference for Augmented Large Language Models

NeurIPS 2025poster

Augmented Large Language Models (LLMs) enhance standalone LLMs by integrating external data sources through API calls. In interactive applications, efficient scheduling is crucial for maintaining low request completion times, directly impacting user engagement. However, these augmentations introduce…

Cited by 11SourceScholar