← Search

André V. Duarte

2 accepted papers

2025

DIS-CO: Discovering Copyrighted Content in VLMs Training Data

ICML 2025poster

*How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data?* Motivated by the hypothesis that a VLM is able to recognize images from its training corpus, we propose DIS-CO, a novel approach to infer the inclusion of…

2024

LumberChunker: Long-Form Narrative Document Segmentation

EMNLP 2024finding

Modern NLP tasks increasingly rely on dense retrieval methods to access up-to-date and relevant contextual information. We are motivated by the premise that retrieval benefits from segments that can vary in size such that a content’s semantic independence is better captured. We propose LumberChunker…