2025
Multitask-Bench: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
COLING 2025main
Recent breakthroughs in Large Language Models (LLMs) have led to their adoption across a wide range of tasks, ranging from code generation to machine translation and sentiment analysis, etc. Red teaming/Safety alignment efforts show that fine-tuning models on benign (non-harmful) data could compromi…