← Search

Lior Belenki

1 accepted papers

2025

Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models

ACL 2025long

We propose a method to optimize language model pre-training data mixtures through efficient approximation of the cross-entropy loss corresponding to each candidate mixture via a Mixture of Data Experts (MDE). We use this approximation as a source of additional features in a regression model, trained…

Cited by 0SourcePDFScholar