← Search

Hui-Liang Shen

22 accepted papers

2026

Large Depth Completion Model from Sparse Observations

ICLR 2026poster

This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without relying on complex architectural designs, LDCM generates metric-accurate dense depth maps in one large transformer. It outpe…

Cited by 0SourceScholar
2026

Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework

AAAI 2026technical

We present MoP-UAV, a new benchmark for UAV-based cross-view object geo-localization guided by multi-modal prompts. MoP-UAV supports fine-grained object-level cross-view localization under diverse prompt modalities, including natural language, bounding boxes, and click points. It offers potential fo

Cited by 0SourcePDFScholar
2026

RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection

CVPR 2026

4D millimeter-wave radar is a promising sensing modality for autonomous driving, yet effective 3D object detection from 4D radar and monocular images remains challenging. Existing fusion approaches either rely on instance proposals lacking global context or dense BEV grids constrained by rigid struc

Cited by 0SourcecodeScholar
2026

Rethinking Unsupervised Cross-modal Flow Estimation: Learning from Decoupled Optimization and Consistency Constraint

ICLR 2026poster

This work presents DCFlow, a novel self-supervised cross-modal flow estimation framework that integrates a decoupled optimization strategy and a cross-modal consistency constraint. Unlike previous unsupervised approaches that implicitly learn flow estimation solely from appearance similarity, we int…

Cited by 0SourceScholar
2025

Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume Construction

ICCV 2025poster

This work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate…

2025

EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image Registration

ICCV 2025poster

Previous deep image registration methods that employ single homography, multi-grid homography, or thin-plate spline often struggle with real scenes containing depth disparities due to their inherent limitations. To address this, we propose an Exponential-Decay Free-Form Deformation Network (EDFFDNet…

Cited by 0SourcePDFScholar
2025

LGDD: Local-Global Synergistic Dual-Branch 3D Object Detection Using 4D Radar

IROS 2025

4D millimeter-wave radar plays a pivotal role in autonomous driving due to its cost-effectiveness and robustness in adverse weather. However, the application of 4D radar point cloud in 3D perception tasks is hindered by its inherent sparsity and noise. To address these challenges, we propose LGDD, a

Cited by 3SourcecodeScholar
2025

Language Driven Occupancy Prediction

ICCV 2025poster

We introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model…

2025

S-BEVLoc: BEV-Based Self-Supervised Framework for Large-Scale LiDAR Global Localization

RA-L 2025

LiDAR-based global localization is an essential component of simultaneous localization and mapping (SLAM), which helps loop closure and re-localization. Current approaches rely on ground-truth poses obtained from GPS or SLAM odometry to supervise network training. Despite the great success of these

Cited by 0SourceScholar
2025

SGDet3D: Semantics and Geometry Fusion for 3D Object Detection Using 4D Radar and Camera

RA-L 2025

4D millimeter-wave radar has gained attention as an emerging sensor for autonomous driving in recent years. However, existing 4D radar and camera fusion models often fail to fully exploit complementary information within each modality and lack deep cross-modal interactions. To address these issues,

Cited by 26SourceScholar
2025

SSHNet: Unsupervised Cross-modal Homography Estimation via Problem Reformulation and Split Optimization

CVPR 2025highlight

We propose a novel unsupervised cross-modal homography estimation learning framework, named Split Supervised Homography estimation Network (SSHNet). SSHNet reformulates the unsupervised cross-modal homography estimation into two supervised sub-problems, each addressed by its specialized network: a h…

2024

Context and Geometry Aware Voxel Transformer for Semantic Scene Completion

NeurIPS 2024spotlight

Vision-based Semantic Scene Completion (SSC) has gained much attention due to its widespread applications in various 3D perception tasks. Existing sparse-to-dense approaches typically employ shared context-independent queries across various input images, which fails to capture distinctions among the…

2024

MCNet: Rethinking the Core Ingredients for Accurate and Efficient Homography Estimation

CVPR 2024poster

We propose Multiscale Correlation searching homography estimation Network namely MCNet an iterative deep homography estimation architecture. Different from previous approaches that achieve iterative refinement by correlation searching within a single scale MCNet combines the multiscale strategy with…

2024

SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning

ECCV 2024poster

"We propose a novel unsupervised cross-modal homography estimation framework based on intra-modal Self-supervised learning, Correlation, and consistent feature map Projection, namely SCPNet. The concept of intra-modal self-supervised learning is first presented to facilitate the unsupervised cross-m…

2023

BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View Images

ICCV 2023poster

Place recognition is a key module for long-term SLAM systems. Current LiDAR-based place recognition methods usually use representations of point clouds such as unordered points or range images. These methods achieve high recall rates of retrieval, but their performance may degrade in the case of vie…

Cited by 46PDFcodeScholar
2023

I2P-Rec: Recognizing Images on Large-Scale Point Cloud Maps Through Bird's Eye View Projections

IROS 2023poster

Place recognition is an important technique for autonomous cars to achieve full autonomy since it can provide an initial guess to online localization algorithms. Although current methods based on images or point clouds have achieved satisfactory performance, localizing the images on a large-scale po…

Cited by 14SourceScholar
2023

Recurrent Homography Estimation Using Homography-Guided Image Warping and Focus Transformer

CVPR 2023poster

We propose the Recurrent homography estimation framework using Homography-guided image Warping and Focus transformer (FocusFormer), named RHWF. Both being appropriately absorbed into the recurrent framework, the homography-guided image warping progressively enhances the feature consistency and the a…

2023

Structure Aggregation for Cross-Spectral Stereo Image Guided Denoising

CVPR 2023poster

To obtain clean images with salient structures from noisy observations, a growing trend in current denoising studies is to seek the help of additional guidance images with high signal-to-noise ratios, which are often acquired in different spectral bands such as near infrared. Although previous guide…

2021

BVMatch: Lidar-Based Place Recognition Using Bird's-Eye View Images

RA-L 2021

Recognizing places using Lidar in large-scale environments is challenging due to the sparse nature of point cloud data. In this letter we present BVMatch, a Lidar-based frame-to-frame place recognition framework, that is capable of estimating 2D relative poses. Based on the assumption that the groun

Cited by 82SourcecodeScholar