Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf’s Law
Recent works have highlighted the optimization difficulties encountered by gradient descent in training the first and last layer of transformer-based language models, which are overcome by optimizers such as Adam. The problem appears linked to the heavy-tailed distribution of words in text data, whe…