2026
QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
ICLR 2026poster
Multimodal Large Language Models (MLLMs) encode images into visual tokens, aligning visual and textual signals within a shared latent space to facilitate cross-modal representation learning. The CLIP model is a widely adopted foundational vision language model whose vision encoder has played a criti…