← Search

Chan Wu

1 accepted papers

2024

Rethinking Memory and Communication Costs for Efficient Data Parallel Training of Large Language Models

NeurIPS 2024poster

Recently, various strategies for distributed training of large language models (LLMs) have been proposed. By categorizing them into basic strategies and composite strategies, we have discovered that existing basic strategies provide limited options in specific scenarios, leaving considerable room fo…

Cited by 0SourcePDFScholar