← Search

Haon Park

3 accepted papers

2026

Jailbreaking on Text-to-Video Models via Scene Splitting Strategy

ICLR 2026poster

Along with the rapid advancement of numerous Text-to-Video (T2V) models, growing concerns have emerged regarding their safety risks. While recent studies have explored vulnerabilities in models like LLMs, VLMs, and Text-to-Image (T2I) models through jailbreak attacks, T2V models remain largely unexp…

Cited by 0SourcecodeScholar
2025

ELITE: Enhanced Language-Image Toxicity Evaluation for Safety

ICML 2025poster

Current Vision Language Models (VLMs) remain vulnerable to malicious prompts that induce harmful outputs. Existing safety benchmarks for VLMs primarily rely on automated evaluation methods, but these methods struggle to detect implicit harmful content or produce inaccurate evaluations. Therefore, we…

Cited by 0SourcePDFScholar
2025

M2S: Multi-turn to Single-turn jailbreak in Red Teaming for LLMs

ACL 2025long

We introduce a novel framework for consolidating multi-turn adversarial “jailbreak” prompts into single-turn queries, significantly reducing the manual overhead required for adversarial testing of large language models (LLMs). While multi-turn human jailbreaks have been shown to yield high attack su…