2026
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
CVPR 2026
Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-specific visual tasks such as medical image diagnosis or fine-grained classification, where representation errors can cascad