2026
Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language
ICML 2026poster
Recent work shows that vision encoders capture ordinal attributes along linear axes, which can be recovered from as few as two labeled images. However, in the zero-shot setting, the text-driven rank axis for Vision-Language Models (VLMs) like CLIP remains suboptimal. In this work, we study the embed…