← Search

Shuangping Li

1 accepted papers

2025

Synthetic continued pretraining

ICLR 2025oral

Pretraining on large-scale, unstructured internet text enables language models to acquire a significant amount of world knowledge. However, this knowledge acquisition is data-inefficient---to learn a fact, models must be trained on hundreds to thousands of diverse representations of it. This poses a…