← Search

Candace Ross

10 accepted papers

2025

DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models

ICCV 2025poster

Recent advances in text-to-image (T2I) models have achieved impressive quality and consistency. However, this has come at the cost of representation diversity. While automatic evaluation methods exist for benchmarking model diversity, they either require reference image datasets or lack specificity…

Cited by 0SourcePDFScholar
2025

Improving Model Evaluation using SMART Filtering of Benchmark Datasets

NAACL 2025long

One of the most challenging problems facing NLP today is evaluation. Some of the most pressing issues pertain to benchmark saturation, data contamination, and diversity in the quality of test examples. To address these concerns, we propose Selection Methodology for Accurate, Reduced, and Targeted (S…

Cited by 2SourcePDFScholar
2025

What’s in Common? Multimodal Models Hallucinate When Reasoning Across Scenes

NeurIPS 2025poster

Multimodal language models possess a remarkable ability to handle an open-vocabulary worth of objects. Yet the best models still suffer from hallucinations when reasoning about scenes in the real world, revealing a gap between their seemingly strong performance on existing perception benchmarks that…

Cited by 0SourceScholar
2024

Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance

ECCV 2024poster

"With the growing popularity of text-to-image generative models, there has been increasing focus on understanding their risks and biases. Recent work has found that state-of-the-art models struggle to depict everyday objects with the true diversity of the real world and have notable gaps between geo…

2024

Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision

AAAI 2024technical

Computer vision models have been known to encode harmful biases, leading to the potentially unfair treatment of historically marginalized groups, such as people of color. However, there remains a lack of datasets balanced along demographic traits that can be used to evaluate the downstream fairness…

2023

FACET: Fairness in Computer Vision Evaluation Benchmark

ICCV 2023poster

Computer vision models have known performance disparities across attributes such as gender and skin tone. This means during tasks such as classification and detection, model performance differs for certain classes based on the demographics of the people in the image. These disparities have been show…

Cited by 46PDFScholar
2022

Perturbation Augmentation for Fairer NLP

EMNLP 2022main

Unwanted and often harmful social biases are becoming ever more salient in NLP research, affecting both models and datasets. In this work, we ask whether training on demographically perturbed data leads to fairer language models. We collect a large dataset of human annotated text perturbations and t…

2022

Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality

CVPR 2022poster

We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly--but crucially, both captions contain a completely identi…

Cited by 440PDFScholar
2020

Learning a natural-language to LTL executable semantic parser for grounded robotics

CoRL 2020

Children acquire their native language with apparent ease by observing how language is used in context and attempting to use it themselves. They do so without laborious annotations, negative examples, or even direct corrections. We take a step toward robots that can do the same by training a grounde

Cited by 0SourcePDFScholar