← Search

Hanwen Sun

1 accepted papers

2026

Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism

ICLR 2026poster

Transformer-based large language models (LLMs) have achieved remarkable success, yet their standard attention mechanism incurs quadratic computation and memory costs with respect to sequence length, posing a major bottleneck for long-context training. Prior work tackles this challenge along two dire…

Cited by 0SourcecodeScholar