Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
Recent research shows that fine-tuning on benign instruction-following data can inadvertently undo the safety alignment process and increase a model's propensity to comply with harmful queries. While instruction-following fine-tuning is important, task-specific fine-tuning-where models are trained o…