Towards Neuro-Symbolic Video Understanding
Minkyu Choi*, Harsh Goel, Mohammad Omama, Yunhao Yang, Sahil Shah, Sandeep Chinchali
Abstract
"The unprecedented surge in video data production in recent years necessitates efficient tools to extract meaningful frames from videos for downstream tasks. Long-term temporal reasoning is a key desideratum for frame retrieval systems. While state-of-the-art foundation models, like VideoLLaMA and ViCLIP, are proficient in short-term semantic understanding, they surprisingly fail at long-term reasoning across frames. A key reason for this failure is that they intertwine per-frame perception and temporal reasoning into a single deep network. Hence, decoupling but co-designing the semantic understanding and temporal reasoning is essential for efficient scene identification. We propose a system that leverages vision-language models for semantic understanding of individual frames and effectively reasons about the long-term evolution of events using state machines and temporal logic (TL) formulae that inherently capture memory. Our TL-based reasoning improves the F1 score of complex event identification by 9 − 15%, compared to benchmarks that use GPT-4 for reasoning, on state-of-the-art self-driving datasets such as Waymo and NuScenes. The source code is available at https://github.com/UTAustin-SwarmLab/Neuro-Symbolic-Video-Search-Temporal-Logic."
BibTeX
@inproceedings{eccv2024_towardsneurosymb,
title = {Towards Neuro-Symbolic Video Understanding},
author = {Minkyu Choi* and Harsh Goel and Mohammad Omama and Yunhao Yang and Sahil Shah and Sandeep Chinchali},
booktitle = {ECCV 2024},
year = {2024}
}