AAAI 2026technical0 citations

Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?

Ming Liu, Hao Chen, Jindong Wang, Liwen Wang, Jingchen Sun, Wensheng Zhang

Abstract

Vision-Language Models (VLMs) have achieved success in tasks such as visual question answering, yet their resilience to distractions remains underexplored. Understanding how distractions affect VLMs

BibTeX
@inproceedings{aaai2026_isyourreasoningm,
  title = {Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?},
  author = {Ming Liu and Hao Chen and Jindong Wang and Liwen Wang and Jingchen Sun and Wensheng Zhang},
  booktitle = {AAAI 2026},
  year = {2026}
}