← Search

Noa Garcia

12 accepted papers

2026

EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories

CVPR 2026

The widespread adoption of text-to-image (T2I) generation has raised concerns about privacy, bias, and copyright violations. Concept erasure techniques offer a promising solution by selectively removing undesired concepts from pre-trained models without requiring full retraining. However, these meth

Cited by 0SourcecodeScholar
2026

ImageSet2Text: Describing Sets of Images Through Text

AAAI 2026technical

In the era of large-scale visual data, understanding collections of images is a challenging yet important task. To this end, we introduce ImageSet2Text, a novel method to automatically generate natural language descriptions of image sets. Based on large language models, visual-question answering cha

Cited by 0SourcePDFScholar
2025

Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation

ICCV 2025poster

Gender bias in vision-language foundation models (VLMs) raises concerns about their safe deployment and is typically evaluated using benchmarks with gender annotations on real-world images. However, as these benchmarks often contain spurious correlations between gender and non-gender features, such…

Cited by 0SourcePDFScholar
2025

Processing and acquisition traces in visual encoders: What does CLIP know about your camera?

ICCV 2025poster

Prior work has analyzed the robustness of visual encoders to image transformations and corruptions, particularly in cases where such alterations are not seen during training. When this occurs, they introduce a form of distribution shift at test time, often leading to performance degradation. The pri…

2024

Can Multiple-choice Questions Really Be Useful in Detecting the Abilities of LLMs?

COLING 2024main

Multiple-choice questions (MCQs) are widely used in the evaluation of large language models (LLMs) due to their simplicity and efficiency. However, there are concerns about whether MCQs can truly measure LLM’s capabilities, particularly in knowledge-intensive scenarios where long-form generation (LF…

2024

Would Deep Generative Models Amplify Bias in Future Models?

CVPR 2024poster

We investigate the impact of deep generative models on potential social biases in upcoming computer vision models. As the internet witnesses an increasing influx of AI-generated images concerns arise regarding inherent biases that may accompany them potentially leading to the dissemination of harmfu…

Cited by 11SourcePDFScholar
2023

Uncurated Image-Text Datasets: Shedding Light on Demographic Bias

CVPR 2023highlight

The increasing tendency to collect large and uncurated datasets to train vision-and-language models has raised concerns about fair representations. It is known that even small but manually annotated datasets, such as MSCOCO, are affected by societal bias. This problem, far from being solved, may be…

2021

Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation

ICCV 2021poster

Have you ever looked at a painting and wondered what is the story behind it? This work presents a framework to bring art closer to people by generating comprehensive descriptions of fine-art paintings. Generating informative descriptions for artworks, however, is extremely challenging, as it require…

Cited by 54PDFcodeScholar
2021

The Met Dataset: Instance-level Recognition for Artworks

NeurIPS 2021poster

This work introduces a dataset for large-scale instance-level recognition in the domain of artworks. The proposed benchmark exhibits a number of different challenges such as large inter-class similarity, long tail distribution, and many classes. We rely on the open access collection of The Met museu…

Cited by 47SourceScholar
2020

Knowledge-Based Video Question Answering with Unsupervised Scene Descriptions

ECCV 2020poster

To understand movies, humans constantly reason over the dialogues and actions shown in specific scenes and relate them to the overall storyline already seen. Inspired by this behaviour, we design ROLL, a model for knowledge-based video story question answering that leverages three crucial aspects of…