Maximizing Intermediate Checkpoint Value in LLM Pretraining with Bayesian Optimization
The rapid proliferation of large language models (LLMs), such as GPT-4 and Gemini, underscores the intense demand for resources during their training processes, posing significant challenges due to substantial computational and environmental costs. In this paper, we introduce a novel checkpoint merg…