2025
JailPO: A Novel Black-Box Jailbreak Framework via Preference Optimization Against Aligned LLMs
AAAI 2025technical
Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipulate prompts to induce harmful outputs. Exploring jailbreak attacks enables us to investigate the vulnerabilities of LLM…