2025
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
ICML 2025spotlight
Vision-language models (VLMs) have improved significantly in their capabilities, but their complex architecture makes their safety alignment challenging. In this paper, we reveal an uneven distribution of harmful information across the intermediate layers of the image encoder and show that skipping…