Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization
Language models are pretrained on sequences that blend statistical regularities (structures making text fluent) with factual associations between specific tokens (corresponding to knowledge of facts). While recent work suggests that the variability of their interaction, such as paraphrases of factua…