2024
Leveraging Web-Crawled Data for High-Quality Fine-Tuning
EMNLP 2024finding
Most large language models are fine-tuned using either expensive human-annotated data or GPT-4 generated data which cannot guarantee performance in certain domains. We argue that although the web-crawled data often has formatting errors causing semantic inaccuracies, it can still serve as a valuable…