← Search

Chengyi Nie

1 accepted papers

2026

A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints

ICML 2026poster

The rapid adoption of large language models (LLMs) has created significant challenges for efficient inference at scale. Unlike traditional workloads, LLM inference is constrained by both computation and the memory overhead of key–value (KV) caching, which accelerates decoding but quickly exhausts GP…

Cited by 0SourceScholar