← Search

Sanyuan Zhao

6 accepted papers

2025

World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous Driving

AAAI 2025technical

The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. However, when faced with perception-limited areas (dynamic or static occlusion regions), MLLMs struggle to effectively integra…

Cited by 2SourcePDFScholar
2022

Modality Synergy Complement Learning with Cascaded Aggregation for Visible-Infrared Person Re-identification

ECCV 2022poster

"Visible-Infrared Re-Identification (VI-ReID) is challenging in image retrievals. The modality discrepancy will easily make huge intra-class variations. Most existing methods either bridge different modalities through modality-invariance or generate the intermediate modality for better performance.…

2021

Cross-Modality Person Re-Identification via Modality Confusion and Center Aggregation

ICCV 2021poster

Cross-modality person re-identification is a challenging task due to large cross-modality discrepancy and intra-modality variations. Currently, most existing methods focus on learning modality-specific or modality-shareable features by using the identity supervision or modality label. Different from…

Cited by 213PDFScholar
2020

Self-Learning With Rectification Strategy for Human Parsing

CVPR 2020poster

In this paper, we solve the sample shortage problem in the human parsing task. We begin with the self-learning strategy, which generates pseudo-labels for unlabeled data to retrain the model. However, directly using noisy pseudo-labels will cause error amplification and accumulation. Considering the…

Cited by 45PDFScholar
2019

Learning Unsupervised Video Object Segmentation Through Visual Attention

CVPR 2019poster

This paper conducts a systematic study on the role of visual attention in Unsupervised Video Object Segmentation (UVOS) tasks. By elaborately annotating three popular video segmentation datasets (DAVIS, Youtube-Objects and SegTrack V2) with dynamic eye-tracking data in the UVOS setting, for the firs…

Cited by 277PDFcodeScholar
2018

Pyramid Dilated Deeper ConvLSTM for Video Salient Object Detection

ECCV 2018poster

This paper proposes a fast video salient object detection model, based on a novel recurrent network architecture, named Pyramid Dilated Bidirectional ConvLSTM (PDB-ConvLSTM). A Pyramid Dilated Convolution (PDC) module is first designed for simultaneously extracting spatial features at multiple scale…

Cited by 593SourcePDFScholar