← Search

Rathul Anand

1 accepted papers

2025

Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures

ICLR 2025poster

Training with larger mini-batches improves the convergence rate and can yield superior performance. However, training with large mini-batches becomes prohibitive for Large Language Models (LLMs), due to the large GPU memory requirement. To address this problem, an effective approach is finding small…

Cited by 0SourcePDFScholar