← Search

Arlindo L. Oliveira

4 accepted papers

2025

DIS-CO: Discovering Copyrighted Content in VLMs Training Data

ICML 2025poster

*How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data?* Motivated by the hypothesis that a VLM is able to recognize images from its training corpus, we propose DIS-CO, a novel approach to infer the inclusion of…

2025

Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

NeurIPS 2025poster

Convolutional neural networks (CNNs) trained on object recognition achieve high task performance but continue to exhibit vulnerability under a range of visual perturbations and out-of-domain images, when compared with biological vision. Prior work has demonstrated that coupling a standard CNN with a…

Cited by 0SourceScholar
2024

DE-COP: Detecting Copyrighted Content in Language Models Training Data

ICML 2024poster

*How can we detect if copyrighted content was used in the training process of a language model, considering that the training data is typically undisclosed?* We are motivated by the premise that a language model is likely to identify verbatim excerpts from its training text. We propose DE-COP, a met…

2024

LumberChunker: Long-Form Narrative Document Segmentation

EMNLP 2024finding

Modern NLP tasks increasingly rely on dense retrieval methods to access up-to-date and relevant contextual information. We are motivated by the premise that retrieval benefits from segments that can vary in size such that a content’s semantic independence is better captured. We propose LumberChunker…