2026
The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
ICML 2026poster
Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability comes with a cost: it can increase a model's tendency to respond to unsafe adversarial prompts, even when fine-tuning w…