2026
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
ICML 2026oral
As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly important to ensure that they are unbiased. In this paper, we argue that the predominant approach of simply removing existing biases from models is not enough. Using a …