← Search

Xiaohan Wang*

1 accepted papers

2024

VideoAgent: Long-form Video Understanding with Large Language Model as Agent

ECCV 2024poster

"Long-form video understanding represents a significant challenge within computer vision, demanding a model capable of reasoning over long multi-modal sequences. Motivated by the human cognitive process for long-form video understanding, we emphasize interactive reasoning and planning over the abili…

Cited by 368SourcePDFScholar