← Search

William Berrios

2 accepted papers

2025

LMUNIT: Fine-grained Evaluation with Natural Language Unit Tests

EMNLP 2025

As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge – human evaluation is costly and noisy, while automated metrics provide only coarse, difficult-to-interpret signals. We introduce natural language unit tests , a paradigm that decompose

2024

Leveraging Diffusion Perturbations for Measuring Fairness in Computer Vision

AAAI 2024technical

Computer vision models have been known to encode harmful biases, leading to the potentially unfair treatment of historically marginalized groups, such as people of color. However, there remains a lack of datasets balanced along demographic traits that can be used to evaluate the downstream fairness…