2026
Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
ICLR 2026poster
Scaling laws for language model training traditionally characterize how performance scales with model size and dataset volume. Prior work has explored architecture variants and data treatments such as dataset filtering and noise injection in language model pretraining; however, these studies have no…