← Search

Zhu Yu

13 accepted papers

2026

Large Depth Completion Model from Sparse Observations

ICLR 2026poster

This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without relying on complex architectural designs, LDCM generates metric-accurate dense depth maps in one large transformer. It outpe…

Cited by 0SourceScholar
2026

Rethinking Unsupervised Cross-modal Flow Estimation: Learning from Decoupled Optimization and Consistency Constraint

ICLR 2026poster

This work presents DCFlow, a novel self-supervised cross-modal flow estimation framework that integrates a decoupled optimization strategy and a cross-modal consistency constraint. Unlike previous unsupervised approaches that implicitly learn flow estimation solely from appearance similarity, we int…

Cited by 0SourceScholar
2025

Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume Construction

ICCV 2025poster

This work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate…

2025

Language Driven Occupancy Prediction

ICCV 2025poster

We introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model…

2025

S-BEVLoc: BEV-Based Self-Supervised Framework for Large-Scale LiDAR Global Localization

RA-L 2025

LiDAR-based global localization is an essential component of simultaneous localization and mapping (SLAM), which helps loop closure and re-localization. Current approaches rely on ground-truth poses obtained from GPS or SLAM odometry to supervise network training. Despite the great success of these

Cited by 0SourceScholar
2025

SGDet3D: Semantics and Geometry Fusion for 3D Object Detection Using 4D Radar and Camera

RA-L 2025

4D millimeter-wave radar has gained attention as an emerging sensor for autonomous driving in recent years. However, existing 4D radar and camera fusion models often fail to fully exploit complementary information within each modality and lack deep cross-modal interactions. To address these issues,

Cited by 26SourceScholar
2025

SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split Optimization

CVPR 2025highlight

We propose a novel unsupervised cross-modal homography estimation learning framework, named Split Supervised Homography estimation Network (SSHNet). SSHNet reformulates the unsupervised cross-modal homography estimation into two supervised sub-problems, each addressed by its specialized network: a h…

2024

Context and Geometry Aware Voxel Transformer for Semantic Scene Completion

NeurIPS 2024spotlight

Vision-based Semantic Scene Completion (SSC) has gained much attention due to its widespread applications in various 3D perception tasks. Existing sparse-to-dense approaches typically employ shared context-independent queries across various input images, which fails to capture distinctions among the…

2023

I2P-Rec: Recognizing Images on Large-Scale Point Cloud Maps Through Bird's Eye View Projections

IROS 2023poster

Place recognition is an important technique for autonomous cars to achieve full autonomy since it can provide an initial guess to online localization algorithms. Although current methods based on images or point clouds have achieved satisfactory performance, localizing the images on a large-scale po…

Cited by 14SourceScholar
2023

Structure Aggregation for Cross-Spectral Stereo Image Guided Denoising

CVPR 2023poster

To obtain clean images with salient structures from noisy observations, a growing trend in current denoising studies is to seek the help of additional guidance images with high signal-to-noise ratios, which are often acquired in different spectral bands such as near infrared. Although previous guide…