← Search

Shuhan Shen

23 accepted papers

2026

BuildingGPT: Auto-Regressive Building Wireframe Reconstruction Model with Reinforcement Learning

CVPR 2026

In this paper, we propose BuildingGPT, a novel auto-regressive model for building wireframe reconstruction from point clouds with reinforcement learning.Unlike prior works based on detection or diffusion models, BuildingGPT reformulates the building wireframe reconstruction task into a sequence pred

Cited by 0SourcecodeScholar
2026

G${2}$VLO: Accurate and Generic 2D Gaussian Based Visual-LiDAR Odometry

RA-L 2026

Multimodal SLAM is an important topic in 3D computer vision research. Recent visual-LiDAR SLAM systems use photometric error for camera pose estimation, but their use of sparse LiDAR projections underutilizes image information. Integrating 3D Gaussian Splatting allows full image rendering, but the e

Cited by 0SourceScholar
2026

PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure Matching

CVPR 2026

While structure-based relocalizers have long strived for point correspondences when establishing or regressing query-map associations, in this paper, we pioneer the use of planar primitives and 3D planar maps for lightweight 6-DoF camera relocalization in structured environments. Planar primitives,

Cited by 0SourcecodeScholar
2026

SRIF: A Safer Ranking Inference Framework for Diffusion Policy Models Without Retraining

RA-L 2026

Recently, diffusion policy models have been applied in the field of robotics. Most existing methods use all observations as condition inputs to the diffusion model. With the denoising of the diffusion model, Gaussian noise gradually becomes an action sequence. However, since these constraints are im

Cited by 0SourceScholar
2025

BWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with Transformer

CVPR 2025highlight

In this paper, we present BWFormer, a novel Transformer-based model for building wireframe reconstruction from airborne LiDAR point cloud. The problem is solved in a ground-up manner here by detecting the building corners in 2D, lifting and connecting them in 3D space afterwards with additional dat…

2025

Causal Enhanced Autoregressive Model for Monocular Image-Goal Navigation in Unknown Map Environment

RA-L 2025

Monocular image-goal navigation in an outdoor environment is a challenging task. Robots have to face monocular scale uncertainty and complex environments. Recently, implementations based on imitation learning have made significant progress. However, robots tend to focus too much on the current state

Cited by 0SourceScholar
2025

CoMatcher: Multi-View Collaborative Feature Matching

CVPR 2025poster

This paper proposes a multi-view collaborative matching strategy for reliable track construction in complex scenarios. We observe that the pairwise matching paradigms applied to image set matching often result in ambiguous estimation when the selected independent pairs exhibit significant occlusions…

Cited by 0SourcePDFScholar
2025

MGSfM: Multi-Camera Geometry Driven Global Structure-from-Motion

ICCV 2025poster

Multi-camera systems are increasingly vital in the environmental perception of autonomous vehicles and robotics. Their physical configuration offers inherent fixed relative pose constraints that benefit Structure-from-Motion (SfM). However, traditional global SfM systems struggle with robustness due…

2025

NeuralPlane: Structured 3D Reconstruction in Planar Primitives with Neural Fields

ICLR 2025oral

3D maps assembled from planar primitives are compact and expressive in representing man-made environments. In this paper, we present **NeuralPlane**, a novel approach that explores **neural** fields for multi-view 3D **plane** reconstruction. Our method is centered upon the core idea of distilling g…

Cited by 0SourcePDFScholar
2025

Quadratic Gaussian Splatting: High Quality Surface Reconstruction with Second-order Geometric Primitives

ICCV 2025poster

We propose Quadratic Gaussian Splatting (QGS), a novel representation that replaces static primitives with deformable quadric surfaces (e.g., ellipse, paraboloids) to capture intricate geometry. Unlike prior works that rely on Euclidean distance for primitive density modeling--a metric misaligned wi…

Cited by 0SourcePDFScholar
2024

BEV2PR: BEV-Enhanced Visual Place Recognition with Structural Cues

IROS 2024

In this paper, we propose a new image-based visual place recognition (VPR) framework by exploiting the structural cues in bird’s-eye view (BEV) from a single monocular camera. The motivation arises from two key observations about place recognition methods based on both appearance and structure: 1) F

Cited by 4SourcecodeScholar
2024

Easing 3D Pattern Reasoning with Side-view Features for Semantic Scene Completion

ECCV 2024poster

"This paper proposes a side-view context inpainting strategy (SidePaint) to ease the reasoning of unknown 3D patterns for semantic scene completion. Based on the observation that the learning burden on pattern completion increases with spatial complexity and feature sparsity, the SidePaint strategy…

Cited by 1SourcePDFScholar
2024

Lightweight Structured Line Map Based Visual Localization

RA-L 2024

Visual localization, also known as camera pose estimation, is a crucial component of many applications, such as robotics, autonomous driving, and augmented reality. Traditional visual localization algorithms typically run on point cloud maps generated by algorithms such as Structure-from-Motion (SfM

Cited by 10SourcecodeScholar
2024

PanoPose: Self-supervised Relative Pose Estimation for Panoramic Images

CVPR 2024highlight

Scaled relative pose estimation i.e. estimating relative rotation and scaled relative translation between two images has always been a major challenge in global Structure-from-Motion (SfM). This difficulty arises because the two-view relative translation computed by traditional geometric vision meth…

Cited by 4SourcePDFScholar
2024

Unsigned Orthogonal Distance Fields: An Accurate Neural Implicit Representation for Diverse 3D Shapes

CVPR 2024poster

Neural implicit representation of geometric shapes has witnessed considerable advancements in recent years. However common distance field based implicit representations specifically signed distance field (SDF) for watertight shapes or unsigned distance field (UDF) for arbitrary shapes routinely suff…

2023

Shape Anchor Guided Holistic Indoor Scene Understanding

ICCV 2023poster

This paper proposes a shape anchor guided learning strategy (AncLearn) for robust holistic indoor scene understanding. We observe that the search space constructed by current methods for proposal feature grouping and instance point sampling often introduces massive noise to instance detection and me…

Cited by 5PDFcodeScholar
2022

Multi-Camera-LiDAR Auto-Calibration by Joint Structure-from-Motion

IROS 2022poster

Multiple sensors, especially cameras and LiDARs, are widely used in autonomous vehicles. In order to fuse data from different sensors accurately, precise calibrations are required, including camera intrinsic parameters, and relative poses between multiple cameras and LiDARs. However, most existing c…

Cited by 20SourceScholar
2021

Semantically Guided Multi-View Stereo for Dense 3D Road Mapping

ICRA 2021poster

Compared to widely used LiDAR-based mapping in autonomous driving field, image-based mapping method has the advantages of low cost, high resolution, and no need for complex calibration. However, the image-based 3D mapping depends heavily on the texture richness and always leaves holes and outliers i…

Cited by 6SourceScholar