2026
BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining
ICML 2026poster
Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream tasks. However, existing approaches often require leveraging external pretrained models, making it difficult to disentangle the effects of data selectio…