← Search

Xiaodong Ji

1 accepted papers

2026

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling

ICML 2026poster

Many advanced Large Language Model (LLM) applications require long-context processing, but the self-attention module becomes a bottleneck during the prefilling stage of inference due to its quadratic time complexity with respect to sequence length. Existing sparse attention methods accelerate attent…

Cited by 0SourcecodeScholar