2026
Segmentation From Attention: Training-Free Layer Selection and One-Shot Tuning for Segmentation in VLMs
ICML 2026poster
Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associations between textual descriptions and image regions. This emergent ability enables zero-shot object detection and segmenta…