2026
Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying Prompts
ICLR 2026poster
Large language models (LLMs) have demonstrated impressive capabilities across a wide range of natural language processing tasks. However, their outputs often exhibit social biases, raising fairness concerns. Existing debiasing methods, such as fine-tuning on additional datasets or prompt engineering…