Revisiting Visual Corruptions in LVLMs: A Shape-Texture Perspective on Model Failures
Large vision-language models (LVLMs) are highly vulnerable to visual corruptions, substantially compromising their reliability and limiting real-world deployment. Prior work has attributed this degradation primarily to insufficient visual grounding and overreliance on language priors. However, these