← Search

Yulin Fei

1 accepted papers

2025

Do Current Video LLMs Have Strong OCR Abilities? A Preliminary Study

COLING 2025main

With the rise of multi-modal large language models, accurately extracting and understanding textual information from video content—referred to as video-based optical character recognition (Video OCR)—has become a crucial capability. This paper introduces a novel benchmark designed to evaluate the vi…