2025
Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge
CVPR 2025poster
Does seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. However, these models frequently encounter a core issue of "cognitive misalignment" between the vision encoder (VE) and the…