← Search

Fengwang Li

1 accepted papers

2025

Understanding and Enhancing the Transferability of Jailbreaking Attacks

ICLR 2025poster

Jailbreaking attacks can effectively manipulate open-source large language models (LLMs) to produce harmful responses. However, these attacks exhibit limited transferability, failing to disrupt proprietary LLMs consistently. To reliably identify vulnerabilities in proprietary LLMs, this work investi…