← Search

Ayoung Lee

2 accepted papers

2026

CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives

ICLR 2026poster

Navigating dilemmas involving conflicting values is challenging even for humans in high-stakes domains, let alone for AI, yet prior work has been limited to everyday scenarios. To close this gap, we introduce CLASH (Character perspective-based LLM Assessments in Situations with High-stakes), a metic…

Cited by 0SourceScholar
2026

LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

ICML 2026poster

Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as lack of exceptionally challenging problems, insufficient test case cove…

Cited by 0SourcecodeScholar