← Search

Ryan Sungmo Kwon

2 accepted papers

2026

CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives

ICLR 2026poster

Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been limited to everyday scenarios. To close this gap, we introduce CLASH (Character perspective-based LLM Assessments in Situations with High-stakes), a metic…

Cited by 0SourceScholar
2025

Representation Bending for Large Language Model Safety

ACL 2025long

Large Language Models (LLMs) have emerged as powerful tools, but their inherent safety risks – ranging from harmful content generation to broader societal harms – pose significant challenges. These risks can be amplified by the recent adversarial attacks, fine-tuning vulnerabilities, and the increas…