← Search

Xuechunzi Bai

2 accepted papers

2026

Large Language Models Develop Novel Social Biases Through Adaptive Exploration

ICML 2026oral

As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In this paper, we argue that the predominant approach of simply removing existing biases from models is not enough. Using a …

Cited by 0SourceScholar
2025

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

ACL 2025long

Although value-aligned language models (LMs) appear unbiased in explicit bias evaluations, they often exhibit stereotypes in implicit word association tasks, raising concerns about their fair usage. We investigate the mechanisms behind this discrepancy and find that alignment surprisingly amplifies…