2025
Bias in the Mirror : Are LLMs opinions robust to their own adversarial attacks
ACL 2025long
Large language models (LLMs) inherit biases from their training data and alignment processes, influencing their responses in subtle ways. While many studies have examined these biases, little work has explored their robustness during interactions. In this paper, we introduce a novel approach where t…