Key Clues Guided Video Character Social Relationship Recognition Enhanced by LLM
Wenlong Dong, Qing Zhu, Qirong Mao
Abstract
Video Character Social Relationship Recognition (VCSRR) requires a comprehensive consideration about spatio-temporal and multi-modal clues in videos. Most existing methods mainly focus on integrating multi-modal clues and modeling interactions among characters. However, they fail to discover key clues in the complex video data or fully understand the clues related to social relationships. In this article, we propose a novel Large Language Model Enhanced Key Clues Selection (LE-KCS) framework to address the aforementioned issues. The core of LE-KCS is to mine multi-scale key clues from the perspectives of time, space and multi-modality, then transfer the knowledge about social relationships of the Large Language Model to VCSRR for understanding the selected clues. We evaluated LE-KCS on the MovieGraphs dataset and the experimental results indicate that our proposed LE-KCS achieves state-of-the-art performance.
BibTeX
@inproceedings{icassp2025_keycluesguidedvi,
title = {Key Clues Guided Video Character Social Relationship Recognition Enhanced by LLM},
author = {Wenlong Dong and Qing Zhu and Qirong Mao},
booktitle = {ICASSP 2025},
year = {2025}
}