← Search

Aishwarya Agarwal

5 accepted papers

2026

Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach

CVPR 2026

Contrastive vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition yet remain vulnerable to spurious correlations, particularly background over-reliance. We introduce Cluster-based Concept Importance (CCI), a novel interpretability method that uses CLIP's own patch embedding

Cited by 0SourceScholar
2025

Composing Parts for Expressive Object Generation

CVPR 2025poster

Image composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, like Stable Diffusion, cannot handle fine-grained part-level attributes in the text prompts. Specifically, when additiona…

Cited by 0SourcePDFScholar
2025

TIDE: Training Locally Interpretable Domain Generalization Models Enables Test-time Correction

CVPR 2025highlight

We consider the problem of single-source domain generalization. Existing methods typically rely on extensive augmentations to synthetically cover diverse domains during training. However, they struggle with semantic shifts (e.g., background and viewpoint changes), as they often learn global features…

Cited by 2SourcePDFScholar
2023

A-STAR: Test-time Attention Segregation and Retention for Text-to-image Synthesis

ICCV 2023poster

While recent developments in text-to-image generative models have led to a suite of high-performing methods capable of producing creative imagery from free-form text, there are several limitations. By analyzing the cross-attention representations of these models, we notice two key issues. First, for…

Cited by 44PDFScholar
2021

MIMOQA: Multimodal Input Multimodal Output Question Answering

NAACL 2021long

Multimodal research has picked up significantly in the space of question answering with the task being extended to visual question answering, charts question answering as well as multimodal input question answering. However, all these explorations produce a unimodal textual output as the answer. In…

Cited by 40SourcePDFScholar