← Search

André Vicente Duarte

1 accepted papers

2024

DE-COP: Detecting Copyrighted Content in Language Models Training Data

ICML 2024poster

*How can we detect if copyrighted content was used in the training process of a language model, considering that the training data is typically undisclosed?* We are motivated by the premise that a language model is likely to identify verbatim excerpts from its training text. We propose DE-COP, a met…