← Search

Hongshan Yu

11 accepted papers

2025

GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector

CVPR 2025poster

We propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is…

2025

InsCMPR: Efficient Cross-Modal Place Recognition via Instance-Aware Hybrid Mamba-Transformer

ICRA 2025

Place recognition is an important technique for autonomous mobile robotic applications. While single-modal sensor-based approaches have shown satisfactory performance, cross-modal place recognition remains underexplored due to the challenge of bridging the cross-modal heterogeneity gap. In this work

Cited by 1SourcecodeScholar
2025

MRMT-PR: A Multi-Scale Reverse-View Mamba-Transformer for LiDAR Place Recognition

IROS 2025

Place recognition is a fundamental technology of high relevance for autonomous robot navigation. Existing methods encounter significant challenges arising from scene variations (e.g., illumination changes, dynamic objects), view-point shifts, and difficulties in data fusion and alignment. These fact

Cited by 1SourceScholar
2025

RID-Net: A Hybrid MLP-Transformer Network for Robust Point Cloud Registration

RA-L 2025

The robustness of correspondence-based point cloud registration relies on transformation invariance and intrinsic distinctiveness of the descriptors computed for registration. However, for challenging scenarios with different objects having similar local geometry and low point cloud overlap, existin

Cited by 0SourceScholar
2024

Fine-Grained Semantic Information Preservation and Misclassification-Aware Loss for 3D Point Cloud

RA-L 2024

Encoder-Decoder structure is a popular choice in point cloud processing for dense multi-classification tasks, e.g., 3D semantic segmentation. Though existing techniques that follow this structure achieve high performance, they are known to suffer from fine-grained information loss, especially when t

Cited by 0SourceScholar
2024

Improved MLP Point Cloud Processing with High-Dimensional Positional Encoding

AAAI 2024technical

Multi-Layer Perceptron (MLP) models are the bedrock of contemporary point cloud processing. However, their complex network architectures obscure the source of their strength. We first develop an “abstraction and refinement” (ABS-REF) view for the neural modeling of point clouds. This view elucidates…

2024

OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition

CVPR 2024poster

Due to the resource-intensive nature of training vision-language models on expansive video data a majority of studies have centered on adapting pre-trained image-language models to the video domain. Dominant pipelines propose to tackle the visual discrepancies with additional temporal learners while…

2024

SGLC: Semantic Graph-Guided Coarse-Fine-Refine Full Loop Closing for LiDAR SLAM

RA-L 2024

Loop closing is a crucial component in SLAM that helps eliminate accumulated errors through two main steps: loop detection and loop pose correction. The first step determines whether loop closing should be performed, while the second estimates the 6-DoF pose to correct odometry drift. Current method

Cited by 12SourcecodeScholar
2023

AShapeFormer: Semantics-Guided Object-Level Active Shape Encoding for 3D Object Detection via Transformers

CVPR 2023poster

3D object detection techniques commonly follow a pipeline that aggregates predicted object central point features to compute candidate points. However, these candidate points contain only positional information, largely ignoring the object-level shape information. This eventually leads to sub-optima…

2023

DANet: Density Adaptive Convolutional Network With Interactive Attention for 3D Point Clouds

RA-L 2023

Local features and contextual dependencies are crucial for 3D point cloud analysis. Many works have been devoted to designing better local convolutional kernels that exploit the contextual dependencies. However, current point convolutions lack robustness to varying point cloud density. Moreover, con

Cited by 7SourceScholar
2022

Learning From Pixel-Level Noisy Label: A New Perspective for Light Field Saliency Detection

CVPR 2022poster

Saliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obta…

Cited by 25PDFcodeScholar