2026
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
ICLR 2026poster
Large language models (LLMs) are increasingly deployed in contexts where their failures have the potential to carry sociopolitical consequences. However, existing safety benchmarks sparsely test vulnerabilities in domains such as political manipulation, propaganda generation, or surveillance and inf…