← Search

Zhipeng Tan

1 accepted papers

2026

DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving

ICLR 2026poster

In large language model (LLM) serving, reusing the key-value (KV) cache of prompts across requests is a key technique for reducing time-to-first-token (TTFT) and lowering serving costs. Cache-affinity scheduling, which co-locates requests with the same prompt prefix to maximize KV cache reuse, often…

Cited by 0SourcecodeScholar