← Search

Jan Wehner

3 accepted papers

2026

Position: Safety Must Precede the Deployment of Open-Ended AI Agents

ICML 2026poster

AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability and adaptability. Within this landscape, open-endedness, where AI agents autonomously and indefinitely generate novel behaviors, representations, or solut…

Cited by 0SourceScholar
2024

Immunization against harmful fine-tuning attacks

EMNLP 2024finding

Large Language Models (LLMs) are often trained with safety guards intended to prevent harmful text generation. However, such safety training can be removed by fine-tuning the LLM on harmful datasets. While this emerging threat (harmful fine-tuning attacks) has been characterized by previous work, th…

Cited by 20SourcePDFScholar
2024

Representation Noising: A Defence Mechanism Against Harmful Finetuning

NeurIPS 2024poster

Releasing open-source large language models (LLMs) presents a dual-use risk since bad actors can easily fine-tune these models for harmful purposes. Even without the open release of weights, weight stealing and fine-tuning APIs make closed models vulnerable to harmful fine-tuning attacks (HFAs). Whi…