2024
On the Relationship between Truth and Political Bias in Language Models
EMNLP 2024main
Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simultaneously can obscure how improving one aspect might impact the others. In this work, we focus on analyzing the relation…