← Search

Robie Gonzales

2 accepted papers

2024

Long-form evaluation of model editing

NAACL 2024long

Evaluations of model editing, a technique for changing the factual knowledge held by Large Language Models (LLMs), currently only use the ‘next few token’ completions after a prompt. As a result, the impact of these methods on longer natural language generation is largely unknown. We introduce long-…

2024

Representation Noising: A Defence Mechanism Against Harmful Finetuning

NeurIPS 2024poster

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). Whi…