← Search

Laiyan Ding

6 accepted papers

2025

DEFOM-Stereo: Depth Foundation Model Based Stereo Matching

CVPR 2025poster

Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization us…

2025

Self-Supervised Enhancement for Depth from a Lightweight ToF Sensor with Monocular Images

IROS 2025

Depth map enhancement using paired high-resolution RGB images offers a cost-effective solution for improving low-resolution depth data from lightweight ToF sensors. Nevertheless, naively adopting a depth estimation pipeline to fuse the two modalities requires groundtruth depth maps for supervision.

Cited by 0SourcecodeScholar
2024

Towards Cross-View-Consistent Self-Supervised Surround Depth Estimation

IROS 2024poster

Depth estimation is a cornerstone for autonomous driving, yet acquiring per-pixel depth ground truth for supervised learning is challenging. Self-Supervised Surround Depth Estimation (SSSDE) from consecutive images offers an economical alternative. While previous SSSDE methods have proposed differen…

Cited by 0SourcecodeScholar
2021

IMENet: Joint 3D Semantic Scene Completion and 2D Semantic Segmentation through Iterative Mutual Enhancement

IJCAI 2021poster

3D semantic scene completion and 2D semantic segmentation are two tightly correlated tasks that are both essential for indoor scene understanding, because they predict the same semantic classes, using positively correlated high-level features. Current methods use 2D features extracted from early-fus…

Cited by 16SourcePDFScholar
2020

DiPE: Deeper into Photometric Errors for Unsupervised Learning of Depth and Ego-motion from Monocular Videos

IROS 2020poster

Unsupervised learning of depth and ego-motion from unlabelled monocular videos has recently drawn great attention, which avoids the use of expensive ground truth in the supervised one. It achieves this by using the photometric errors between the target view and the synthesized views from its adjacen…

Cited by 24SourcecodeScholar