← Search

Brendan Murphy

3 accepted papers

2025

Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility

EMNLP 2025

AI systems are rapidly advancing in capability, and frontier model developers broadly acknowledge the need for safeguards against serious misuse. However, this paper demonstrates that fine-tuning, whether via open weights or closed fine-tuning APIs, can produce helpful-only models with safeguards de

2025

On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

ICLR 2025poster

As LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators. However, training to maximize human feedback creates a perverse incentive structure for the AI to resort to manipulative…

2025

Scaling Trends for Data Poisoning in LLMs

AAAI 2025technical

LLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-tuning on poisoned data, even when safeguarded by moderation systems. Given the persistence of data poisoning vulnerabili…