← Search

Tianming Liang

7 accepted papers

2026

Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible

Cited by 0SourcecodeScholar
2026

Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation

CVPR 2026

Referring Video Object Segmentation (RVOS) aims to segment objects in videos based on textual queries. Current methods mainly rely on large-scale supervised fine-tuning (SFT) of Multi-modal Large Language Models (MLLMs). However, this paradigm suffers from heavy data dependence and limited scalabili

Cited by 0SourcecodeScholar
2026

Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search

ICML 2026poster

Segmentation based on language has been a popular topic in computer vision. While recent advances in multimodal large language models (MLLMs) have endowed segmentation systems with reasoning capabilities, these efforts remain confined by the frozen internal knowledge of MLLMs, which limits their pot…

Cited by 0SourceScholar
2026

TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding

AAAI 2026technical

Spatio-Temporal Video Grounding (STVG) aims to localize a spatio-temporal tube that corresponds to a given language query in an untrimmed video. This is a challenging task since it involves complex vision-language understanding and spatiotemporal reasoning. Recent works have explored weakly-superv

Cited by 0SourcePDFScholar
2025

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

ICCV 2025poster

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existi…

Cited by 0SourcePDFScholar
2024

Progressive Pretext Task Learning for Human Trajectory Prediction

ECCV 2024poster

"Human trajectory prediction is a practical task of predicting the future positions of pedestrians on the road, which typically covers all temporal ranges from short-term to long-term within a trajectory. However, existing works attempt to address the entire trajectory prediction with a singular, un…

2024

Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels

CVPR 2024poster

This paper focuses on open-ended video question answering which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task since a question may have multiple answers. However due to annotation costs the labe…

Cited by 2SourcePDFScholar