2024
Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning
ACL 2024findings
Through pretraining on a corpus with various sources, Large Language Models (LLMs) have gained impressive performance. However, the impact of each component of the pretraining corpus remains opaque. As a result, the organization of the pretraining corpus is still empirical and may deviate from the o…