2026
Silenced Biases: The Dark Side LLMs Learned to Refuse
AAAI 2026technical
Safety-aligned large language models (LLMs) are becoming increasingly widespread, especially in sensitive applications where fairness is essential and biased outputs can cause significant harm. However, evaluating the fairness of models is a complex challenge, and approaches that do so typically uti