← Search

Xiaoning Dong

2 accepted papers

2025

SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage

ACL 2025finding

Large language models (LLMs) have made significant advancements across various tasks, but their safety alignment remains a major concern. Exploring jailbreak prompts can expose LLMs’ vulnerabilities and guide efforts to secure them. Existing methods primarily design sophisticated instructions for th…

2025

SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language Models

EMNLP 2025

Multimodal large language models (MLLMs) demonstrate impressive capabilities by integrating visual and textual information. However, the incorporation of visual modalities also introduces new and complex safety risks, rendering even the most advanced models vulnerable to sophisticated jailbreak atta