2025
Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation
ACL 2025long
Standard bias benchmarks used for large language models (LLMs) measure the association between social attributes in model inputs and single-word model outputs. We test whether these benchmarks are robust to lengthening the model outputs via a more realistic user prompt, in the commonly studied domai…