2025
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
ACL 2025finding
Image captioning has been a longstanding challenge in vision-language research. With the rise of LLMs, modern Vision-Language Models (VLMs) generate detailed and comprehensive image descriptions. However, benchmarking the quality of such captions remains unresolved. This paper addresses two key ques…