2025
Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study
COLING 2025main
With the rise of multi-modal large language models, accurately extracting and understanding textual information from video content—referred to as video-based optical character recognition (Video OCR)—has become a crucial capability. This paper introduces a novel benchmark designed to evaluate the vi…