← Search

Xiaojian Ma*

1 accepted papers

2024

VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

ECCV 2024poster

"We explore how reconciling several foundation models (large language models and vision-language models) with a novel unified memory mechanism could tackle the challenging video understanding problem, especially capturing the long-term temporal relations in lengthy videos. In particular, the propose…