AAAI 2026technical0 citations

Silenced Biases: The Dark Side LLMs Learned to Refuse

Rom Himelstein, Amit LeVi, Brit Youngmann, Yaniv Nemcovsky, Avi Mendelson

Abstract

Safety-aligned large language models (LLMs) are becoming increasingly widespread, especially in sensitive applications where fairness is essential and biased outputs can cause significant harm. However, evaluating the fairness of models is a complex challenge, and approaches that do so typically utilize standard question-answer (QA) styled schemes. Such methods often overlook deeper issues by interpreting the model

BibTeX
@inproceedings{aaai2026_silencedbiasesth,
  title = {Silenced Biases: The Dark Side LLMs Learned to Refuse},
  author = {Rom Himelstein and Amit LeVi and Brit Youngmann and Yaniv Nemcovsky and Avi Mendelson},
  booktitle = {AAAI 2026},
  year = {2026}
}
Silenced Biases: The Dark Side LLMs Learned to Refuse · AAAI 2026