2026
Training–Inference Consistent Segmented Execution for Long-Context LLMs
ICML 2026poster
Transformer-based large language models face severe scalability challenges in long-context generation due to the computational and memory costs of full-context attention. Under practical computation and memory constraints, many inference-efficient long-context methods improve efficiency by adopting …