2025
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
ICLR 2025poster
Training with larger mini-batches improves the convergence rate and can yield superior performance. However, training with large mini-batches becomes prohibitive for Large Language Models (LLMs), due to the large GPU memory requirement. To address this problem, an effective approach is finding small…