← Search

Weichen Xiong

1 accepted papers

2024

SoftDedup: an Efficient Data Reweighting Method for Speeding Up Language Model Pre-training

ACL 2024long

The effectiveness of large language models (LLMs) is often hindered by duplicated data in their extensive pre-training datasets. Current approaches primarily focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. To a…

Cited by 2SourcePDFScholar