External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video Understanding
Long video understanding with Large Language Models (LLMs) enables the description of objects that are not explicitly present in the training data. However, continuous changes in known objects and the emergence of new ones require up-to-date knowledge of objects and their dynamics for effective unde