2024
Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
NeurIPS 2024poster
Safety alignment is crucial to ensure that large language models (LLMs) behave in ways that align with human preferences and prevent harmful actions during inference. However, recent studies show that the alignment can be easily compromised through finetuning with only a few adversarially designed t…