← Search

Yanhao Dong

1 accepted papers

2026

Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching

AAAI 2026technical

Large Language Models (LLMs) exhibit pronounced memory-bound characteristics during inference due to High Bandwidth Memory (HBM) bandwidth constraints. In this paper, we propose an L2 Cache-oriented asynchronous KV Cache prefetching method to break through the memory bandwidth bottleneck in LLM infe

Cited by 0SourcePDFScholar