← Search

Seonghwan Choi

1 accepted papers

2026

Retrospective Sparse Attention for Efficient Long-Context Generation

ICLR 2026poster

Large Language Models (LLMs) are increasingly deployed in long-context tasks such as reasoning, code generation, and multi-turn dialogue. However, inference over extended contexts is bottlenecked by the Key-Value (KV) cache, whose memory footprint grows linearly with sequence length and dominates la…

Cited by 0SourcecodeScholar