2026
Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
ICML 2026poster
Deep neural networks with repeated blocks, such as transformers and ResNets, often exhibit closely related representational structure across layers that emerges with training. Motivated by this observation, we introduce *Gradient Smoothing*, a general training paradigm that couples gradient updates …