2025
Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
ACL 2025finding
The safety alignment ability of Vision-Language Models (VLMs) is prone to be degraded by the integration of the vision module compared to its LLM backbone. We investigate this phenomenon, dubbed as “safety alignment degradation” in this paper, and show that the challenge arises from the representati…