← Search

Adriana Kovashka

22 accepted papers

2026

Culture in Action: Evaluating Text-to-Image Models through Social Activities

ICLR 2026poster

Text-to-image (T2I) diffusion models achieve impressive photorealism by training on large-scale web data, but models inherit cultural biases and fail to depict underrepresented regions faithfully. Existing cultural benchmarks focus mainly on object-centric categories (e.g., food, attire, and archit…

Cited by 0SourceScholar
2025

Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition

ICASSP 2025accepted

First-person activity recognition is rapidly growing due to the widespread use of wearable cameras but faces challenges from domain shifts across different environments, such as varying objects or background scenes. We propose a multi-modal framework that improves domain generalization by integratin…

Cited by 0SourceScholar
2025

Leveraging Large Models to Evaluate Novel Content: A Case Study on Advertisement Creativity

EMNLP 2025

Evaluating creativity is challenging, even for humans, not only because of its subjectivity but also because it involves complex cognitive processes. Inspired by work in marketing, we attempt to break down visual advertisement creativity into atypicality and originality. With fine-grained human anno

2025

Role Bias in Diffusion Models: Diagnosing and Mitigating through Intermediate Decomposition

NeurIPS 2025poster

Text-to-image (T2I) diffusion models exhibit impressive photorealistic image generation capabilities, yet they struggle in compositional image generation. In this work, we introduce RoleBench, a benchmark focused on evaluating compositional generalization in action-based relations (e.g., "mouse chas…

Cited by 0SourceScholar
2024

Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition

CVPR 2024poster

Existing object recognition models have been shown to lack robustness in diverse geographical scenarios due to domain shifts in design and context. Class representations need to be adapted to more accurately reflect an object concept under these shifts. In the absence of training data from target ge…

Cited by 3SourcePDFScholar
2024

Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval

EMNLP 2024main

There is a scarcity of multilingual vision-language models that properly account for the perceptual differences that are reflected in image captions across languages and cultures. In this work, through a multimodal, multilingual retrieval case study, we quantify the existing lack of model flexibilit…

Cited by 1SourcePDFScholar
2024

Synonym relations affect object detection learned on vision-language data

NAACL 2024findings

We analyze whether object detectors trained on vision-language data learn effective visual representations for synonyms. Since many current vision-language models accept user-provided textual input, we highlight the need for such models to learn feature representations that are robust to changes in…

2021

Domain-Robust VQA With Diverse Datasets and Methods but No Target Labels

CVPR 2021poster

The observation that computer vision methods overfit to dataset specifics has inspired diverse attempts to make object recognition models robust to domain shifts. However, similar work on domain-robust visual question answering methods is very limited. Domain adaptation for VQA differs from adaptati…

Cited by 29PDFScholar
2019

Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection

ICCV 2019poster

Learning to localize and name object instances is a fundamental problem in vision, but state-of-the-art approaches rely on expensive bounding box supervision. While weakly supervised detection (WSOD) methods relax the need for boxes to that of image-level annotations, even cheaper supervision is nat…

Cited by 61PDFcodeScholar
2017

Automatic Understanding of Image and Video Advertisements

CVPR 2017spotlight

There is more to images than their objective physical content: for example, advertisements are created to persuade a viewer to take a certain action. We propose the novel problem of automatic advertisement understanding. To enable research on this problem, we create two datasets: an image dataset of…

Cited by 219PDFScholar