2026
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking
ICML 2026poster
Efficiently processing long sequences with Transformer models usually requires splitting the computations across accelerators via context parallelism. The dominant approaches in this family of methods, such as Ring Attention or DeepSpeed Ulysses, enable scaling over the context dimension but do not …