← Search

Sho Takishita

1 accepted papers

2025

LLMs Can Compensate for Deficiencies in Visual Representations

EMNLP 2025

Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We investigate the hypothesis that the strong language backbone in VLMs compensates for possibly weak visual features by conte

Cited by 0SourcePDFScholar