2025
Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
ACL 2025long
Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs).However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous creation of key points, exorbitant cost of data creation, and li…