← Search

Zichuan Xu

6 accepted papers

2024

Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language

AAAI 2024technical

Given an untrimmed video and a sentence query, video moment retrieval using language (VMR) aims to locate a target query-relevant moment. Since the untrimmed video is overlong, almost all existing VMR methods first sparsely down-sample each untrimmed video into multiple fixed-length video clips and…

Cited by 19SourcePDFScholar
2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding

AAAI 2022technical

Temporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although existing methods train well-designed deep networks with large amount of data, we find that they can easily forget the rarely appeared cases during training due to the off-balance data distribution, which i…

Cited by 66SourcePDFScholar
2022

Rethinking the Video Sampling and Reasoning Strategies for Temporal Sentence Grounding

EMNLP 2022finding

Temporal sentence grounding (TSG) aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. All existing works first utilize a sparse sampling strategy to extract a fixed number of video frames and then interact them with query for reasoning.However, w…

Cited by 22SourcePDFScholar
2022

Unsupervised Temporal Video Grounding with Deep Semantic Clustering

AAAI 2022technical

Temporal video grounding (TVG) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant video-query paired data, which is expensive to collect in real-world scenarios. In this…

Cited by 59SourcePDFScholar
2021

Context-Aware Biaffine Localizing Network for Temporal Sentence Grounding

CVPR 2021poster

This paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. Previous works either compare pre-defined candidate segments with the query and select the best one by ranking, or di…

Cited by 176PDFcodeScholar
2021

Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object Segmentation

AAAI 2021technical

This paper addresses the task of segmenting class-agnostic objects in semi-supervised setting. Although previous detection based methods achieve relatively good performance, these approaches extract the best proposal by a greedy strategy, which may lose the local patch details outside the chosen can…

Cited by 27SourcePDFScholar