← Search

Samuele Poppi

1 accepted papers

2025

Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks

NAACL 2025findings

Recent advancements in Large Language Models (LLMs) have sparked widespread concerns about their safety. Recent work demonstrates that safety alignment of LLMs can be easily removed by fine-tuning with a few adversarially chosen instruction-following examples, i.e., fine-tuning attacks. We take a fu…

Cited by 8SourcePDFScholar