Difficulty–Diversity Collaborative Filtering for Data-Efficient LLM Fine-Tuning
The performance of fine-tuned language models is heavily influenced by the quality and quantity of their fine-tuning data. While scaling laws suggest that larger models benefit from more data during pretraining, the Less-is-More hypothesis highlights that downstream fine-tuning often requires only a…