2025
Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian Framework
NeurIPS 2025poster
Careful curation of data sources can significantly improve the performance of LLM pre-training, but predominant approaches rely heavily on intuition or costly trial-and-error, making them difficult to generalize across different data domains and downstream tasks. Although scaling laws can provide a…