← Search

Yongjian Zhang

7 accepted papers

2026

Seeing Motion, Generating Action: Explicit Motion-Aware Policy for Robotic Action Generation

ICRA 2026poster

Imitation learning (IL) offers a scalable framework for teaching robots complex manipulation skills from human demonstrations. However, conventional end-to-end visuomotor IL models often suffer from poor performance and robustness due to the significant modality mismatch between high-dimensional vis…

Cited by 0Scholar
2025

DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label Supervision

AAAI 2025technical

Self-supervised stereo matching has drawn attention due to its ability to estimate disparity without needing ground-truth data. However, existing self-supervised stereo matching methods heavily rely on the photo-metric consistency assumption, which is vulnerable to natural disturbances, resulting in…

Cited by 0SourcePDFScholar
2025

Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts

ICCV 2025poster

Recently, learning-based stereo matching networks have advanced significantly.However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets.Leveraging Vision Foundation Models (VFMs)…

2025

PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo Matching

NeurIPS 2025poster

Temporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersion of users. Despite its importance, this task remains challenging due to the difficulty in modeling long-term temporal…

Cited by 0SourcecodeScholar
2025

Self-Distilled Stereo Matching: Real-Time Domain Generalization for Robotic Depth Perception

IROS 2025

While human vision inherently achieves robust cross-domain depth estimation through binocular coordination, robotic systems employing stereo matching still confront significant challenges in maintaining robustness across domains when performing real-time environmental depth perception. Furthermore,

Cited by 0SourceScholar
2024

Learning Representations from Foundation Models for Domain Generalized Stereo Matching

ECCV 2024poster

"State-of-the-art stereo matching networks trained on in-domain data often underperform on cross-domain scenes. Intuitively, leveraging the zero-shot capacity of a foundation model can alleviate the cross-domain generalization problem. The main challenge of incorporating a foundation model into ster…

Cited by 7SourcePDFScholar
2022

SLFNet: A Stereo and LiDAR Fusion Network for Depth Completion

RA-L 2022

Acquiring dense and precise depth information in real time is highly demanded for robotic perception and automatic driving. Motivated by the complementary nature of stereo images and LiDAR point clouds, we propose an efficient stereo-LiDAR fusion network (SLFNet) to predict a dense depth map of a sc

Cited by 14SourceScholar