2025
Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
ICLR 2025poster
In this work, we investigate whether small language models can determine high-quality subsets of large-scale text datasets that improve the performance of larger language models. While existing work has shown that pruning based on the perplexity of a larger model can yield high-quality data, we inve…