2026
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
ICML 2026poster
Large-scale web-crawled datasets contain noise, bias, and irrelevant information, necessitating data selection techniques. Existing methods depend on hand-crafted heuristics, downstream datasets, or require expensive influence-based computations---all of which limit scalability and introduce unwante…