← Search

Xianpeng Shang

1 accepted papers

2026

Training–Inference Consistent Segmented Execution for Long-Context LLMs

ICML 2026poster

Transformer-based large language models face severe scalability challenges in long-context generation due to the computational and memory costs of full-context attention. Under practical computation and memory constraints, many inference-efficient long-context methods improve efficiency by adopting …

Cited by 0SourceScholar