2024
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
ECCV 2024poster
"We explore how reconciling several foundation models (large language models and vision-language models) with a novel unified memory mechanism could tackle the challenging video understanding problem, especially capturing the long-term temporal relations in lengthy videos. In particular, the propose…