← Search

Seorin Kim

1 accepted papers

2025

KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models

EMNLP 2025

Large language models (LLMs) often exhibit societal biases in their outputs, prompting ethical concerns regarding fairness and harm. In this work, we propose KLAAD (KL-Attention Alignment Debiasing), an attention-based debiasing framework that implicitly aligns attention distributions between stereo

Cited by 0SourcePDFScholar