Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues
Training large neural networks exposes neural scaling laws for the generalization error, which points to a universal behavior across network architectures of learning in high dimensions. It was also shown that this effect persists in the limit of highly overparametrized networks as well as the Neura…