2025
IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
COLING 2025main
Large-scale Vision-Language Models (LVLMs) process both images and text, excelling in multimodal tasks such as image captioning and description generation. However, while these models excel at generating factual content, their ability to generate and evaluate texts reflecting perspectives on the sam…