← Search

Chaitanya Baranwal

1 accepted papers

2023

Sequence Parallelism: Long Sequence Training from System Perspective

ACL 2023long

Transformer achieves promising results on various tasks. However, self-attention suffers from quadratic memory requirements with respect to the sequence length. Existing work focuses on reducing time and space complexity from an algorithm perspective. In this work, we propose sequence parallelism, a…

Cited by 102SourcePDFScholar