AAAI 2026technical0 citations
TVChain: Leveraging Textual-Visual Prompt Chains for Jailbreaking Large Vision-Language Models
Hao Yu, Ke Liang, Junxian Duan, Jun Wang, Siwei Wang, Chuan Ma, Xinwang Liu
Abstract
Large Vision-Language Models (LVLMs) enhance the capabilities of Large Language Models by integrating visual inputs, thereby enabling advanced multimodal reasoning across diverse applications. However, these enhanced reasoning capabilities introduce new security risks, particularly to jailbreaking attacks that bypass built-in safety mechanisms to elicit harmful or unauthorized outputs. While recent efforts have explored adversarial and typographic prompts, most existing attacks suffer from three key limitations: reliance on auxiliary models, limited effectiveness in black-box scenarios, and inadequate exploitation of the LVLMs
BibTeX
@inproceedings{aaai2026_tvchainleveragin,
title = {TVChain: Leveraging Textual-Visual Prompt Chains for Jailbreaking Large Vision-Language Models},
author = {Hao Yu and Ke Liang and Junxian Duan and Jun Wang and Siwei Wang and Chuan Ma and Xinwang Liu},
booktitle = {AAAI 2026},
year = {2026}
}