← Search

Alexander Nicholas D’Amour

1 accepted papers

2025

Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation

ACL 2025long

Standard bias benchmarks used for large language models (LLMs) measure the association between social attributes in model inputs and single-word model outputs. We test whether these benchmarks are robust to lengthening the model outputs via a more realistic user prompt, in the commonly studied domai…

Cited by 0SourcePDFScholar