← Search

Matthew Daniel Hull

1 accepted papers

2024

Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models

NeurIPS 2024poster

Safety alignment is crucial to ensure that large language models (LLMs) behave in ways that align with human preferences and prevent harmful actions during inference. However, recent studies show that the alignment can be easily compromised through finetuning with only a few adversarially designed t…