2025
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
ACL 2025finding
Despite a growing literature finding that large language models (LLMs) exhibit demographic biases, reports with whom they align best are hard to generalize or even contradictory. In this work, we examine the alignment of LLMs with human annotations in five offensive language datasets, comprising app…