2026
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
ICLR 2026poster
Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex referential language and diminish both the accuracy and effic…