2025
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
NAACL 2025findings
Recent advancements in Large Language Models (LLMs) have sparked widespread concerns about their safety. Recent work demonstrates that safety alignment of LLMs can be easily removed by fine-tuning with a few adversarially chosen instruction-following examples, i.e., fine-tuning attacks. We take a fu…