A universal compression theory: Lottery ticket hypothesis and superpolynomial scaling laws
When training large-scale models, the performance typically scales with the number of parameters and the dataset size according to a slow power law. A fundamental theoretical and practical question is whether comparable performance can be achieved with significantly smaller models and substantially…