← Search

Lê Hải Sơn

1 accepted papers

2026

SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests

ICLR 2026poster

Large language models (LLMs) are increasingly deployed in contexts where their failures have the potential to carry sociopolitical consequences. However, existing safety benchmarks sparsely test vulnerabilities in domains such as political manipulation, propaganda generation, or surveillance and inf…

Cited by 0SourcecodeScholar