2025
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling
NeurIPS 2025poster
Modern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers,…