← Search

Chaolei Tan

9 accepted papers

2026

Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible

Cited by 0SourcecodeScholar
2026

TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding

AAAI 2026technical

Spatio-Temporal Video Grounding (STVG) aims to localize a spatio-temporal tube that corresponds to a given language query in an untrimmed video. This is a challenging task since it involves complex vision-language understanding and spatiotemporal reasoning. Recent works have explored weakly-superv

Cited by 0SourcePDFScholar
2025

Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

ICCV 2025poster

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos…

2025

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

ICCV 2025poster

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existi…

Cited by 0SourcePDFScholar
2025

SAUGE: Taming SAM for Uncertainty-Aligned Multi-Granularity Edge Detection

AAAI 2025technical

Edge labels are typically at various granularity levels owing to the varying preferences of annotators, thus handling the subjectivity of per-pixel labels has been a focal point for edge detection. Previous methods often employ a simple voting strategy to diminish such label uncertainty or impose a…

2024

Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels

CVPR 2024poster

This paper focuses on open-ended video question answering which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task since a question may have multiple answers. However due to annotation costs the labe…

Cited by 2SourcePDFScholar
2024

Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding

CVPR 2024poster

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However existing VPG approaches are heavily reliant on a considerable number of temporal labels that are…

Cited by 4SourcePDFScholar
2023

Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding

CVPR 2023poster

Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the la…

2023

Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding

CVPR 2023poster

Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex sem…

Cited by 24SourcePDFScholar