A STUDY OF DATA SELECTION STRATEGIES FOR PRE-TRAINING SELF-SUPERVISED SPEECH MODELS
Self-supervised learning (SSL) has transformed speech processing, yet its reliance on massive pre-training datasets remains a bottleneck. While robustness is often attributed to scale and diversity, the role of the data distribution is less understood. We systematically examine how curated subsets o…