2024
DE-COP: Detecting Copyrighted Content in Language Models Training Data
ICML 2024poster
*How can we detect if copyrighted content was used in the training process of a language model, considering that the training data is typically undisclosed?* We are motivated by the premise that a language model is likely to identify verbatim excerpts from its training text. We propose DE-COP, a met…