Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
Reinforcement learning (RL), particularly RL from verifiable reward (RLVR), has become a crucial phase of training large language models (LLMs) and a key focus of current scaling efforts. However, optimization practices in RL largely follow those of next-token-prediction stages (e.g., pretraining an…