← Search

Sangyoon Yu

3 accepted papers

2025

ELITE: Enhanced Language-Image Toxicity Evaluation for Safety

ICML 2025poster

Current Vision Language Models (VLMs) remain vulnerable to malicious prompts that induce harmful outputs. Existing safety benchmarks for VLMs primarily rely on automated evaluation methods, but these methods struggle to detect implicit harmful content or produce inaccurate evaluations. Therefore, we…

Cited by 0SourcePDFScholar
2025

M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs

ACL 2025long

We introduce a novel framework for consolidating multi-turn adversarial “jailbreak” prompts into single-turn queries, significantly reducing the manual overhead required for adversarial testing of large language models (LLMs). While multi-turn human jailbreaks have been shown to yield high attack su…

2023

DepthFL : Depthwise Federated Learning for Heterogeneous Clients

ICLR 2023poster

Federated learning is for training a global model without collecting private local data from clients. As they repeatedly need to upload locally-updated weights or gradients instead, clients require both computation and communication resources enough to participate in learning, but in reality their r…