2025
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
ACL 2025long
Although value-aligned language models (LMs) appear unbiased in explicit bias evaluations, they often exhibit stereotypes in implicit word association tasks, raising concerns about their fair usage. We investigate the mechanisms behind this discrepancy and find that alignment surprisingly amplifies…