← Search

Kai Williams

2 accepted papers

2024

Immunization against harmful fine-tuning attacks

EMNLP 2024finding

Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM on harmful datasets. While this emerging threat (harmful fine-tuning attacks) has been characterized by previous work, th…

Cited by 20SourcePDFScholar
2024

Representation Noising: A Defence Mechanism Against Harmful Finetuning

NeurIPS 2024poster

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). Whi…