AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens. However, existing LLM training libraries do not provide easy to use abstractions to optimize for long-context training, instead fo…