ICASSP 2024accepted0 citations

Mrtnet: Multi-Resolution Temporal Network for Video Sentence Grounding

Wei Ji, You Qin, Long Chen, Yinwei Wei, Yiming Wu, Roger Zimmermann

Abstract

Video sentence grounding locates a specific moment in a video based on a text query. Existing methods focus on single temporal resolution, ignoring multi-scale temporal consistency. We introduce MRTNet, a multi-resolution grounding network with four key components: a feature encoder, a Multi-Resolution Temporal (MRT) module, a Query-aware Attention (QAM) module, and a predictor. The MRT module uses an encoder-decoder network and Transformers to predict start and end times. The QAM module fuses visual and text features. Both MRT and QAM modules are easily integrated into existing VSG models. We also employ a loss function for cross-modal feature supervision at multiple scales. Extensive experiments on two prevalent datasets have shown the effectiveness of MRTNet.

BibTeX
@inproceedings{icassp2024_mrtnetmultiresol,
  title = {Mrtnet: Multi-Resolution Temporal Network for Video Sentence Grounding},
  author = {Wei Ji and You Qin and Long Chen and Yinwei Wei and Yiming Wu and Roger Zimmermann},
  booktitle = {ICASSP 2024},
  year = {2024}
}
Mrtnet: Multi-Resolution Temporal Network for Video Sentence Grounding · ICASSP 2024