2025
AutoClean: LLMs Can Prepare Their Training Corpus
NAACL 2025system demonstrations
Recent studies highlight the reliance of Large Language Models (LLMs) on high-quality, diverse data for optimal performance. The data sourced from the Internet often aggregated into datasets like the Common Crawl corpus, presents significant quality variability and necessitates extensive cleaning. M…