2025
ZeCO: Zero-Communication Overhead Sequence Parallelism for Linear Attention
NeurIPS 2025poster
Linear attention mechanisms deliver significant advantages for Large Language Models (LLMs) by providing linear computational complexity, enabling efficient processing of ultra-long sequences (e.g., 1M context). However, existing Sequence Parallelism (SP) methods, essential for distributing these wo…