2025
Breach By A Thousand Leaks: Unsafe Information Leakage in 'Safe' AI Responses
ICLR 2025poster
Vulnerability of Frontier language models to misuse has prompted the development of safety measures like filters and alignment training seeking to ensure safety through robustness to adversarially crafted prompts. We assert that robustness is fundamentally insufficient for ensuring safety goals due…