2025
STAMP Your Content: Proving Dataset Membership via Watermarked Rephrasings
ICML 2025poster
Given how large parts of publicly available text are crawled to pretrain large language models (LLMs), data creators increasingly worry about the inclusion of their proprietary data for model training without attribution or licensing. Their concerns are also shared by benchmark curators whose test-s…