← Search

Yingsheng Geng

1 accepted papers

2026

RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse

ICML 2026poster

The increasing complexity of AI tasks has shifted the paradigm from monolithic models toward multi-agent large language model (LLM) systems. However, these collaborative architectures introduce a critical bottleneck: redundant prefill computation for shared content generated by previous agents, whic…

Cited by 0SourceScholar