CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Extensive training datasets represent one of the important factors for the impressive learning capabilities of large language models (LLMs). However, these training datasets for current LLMs, especially the recent state-of-the-art models, are often not fully disclosed. Creating training data for hig…