2025
DIS-CO: Discovering Copyrighted Content in VLMs Training Data
ICML 2025poster
*How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data?* Motivated by the hypothesis that a VLM is able to recognize images from its training corpus, we propose DIS-CO, a novel approach to infer the inclusion of…