← Search

Hongyuan Yuan

4 accepted papers

2026

Asking like Socrates: Socrates helps VLMs understand remote sensing images

CVPR 2026

Recent multimodal reasoning models, inspired by DeepSeek-R1, have significantly advanced vision-language systems. However, in remote sensing (RS) tasks, we observe widespread pseudo reasoning: models narrate the process of reasoning rather than genuinely reason toward the correct answer based on vis

Cited by 0SourcecodeScholar
2024

A Hybrid Approach for Cross-Modality Pose Estimation Between Image and Point Cloud

RA-L 2024

Cross-modality pose estimation/localization is a critical challenge for multi-sensor-based perception systems, with applications spanning vehicle localization and online calibrations. In this paper, we introduce a hybrid approach to estimate the camera pose with respect to a point cloud with co-visi

Cited by 1SourceScholar
2024

NDT-Map-Code: A 3D global descriptor for real-time loop closure detection in lidar SLAM

IROS 2024poster

Loop-closure detection, also known as place recognition, aiming to identify previously visited locations, is an essential component of a SLAM system. Existing research on lidar-based loop closure heavily relies on dense point cloud and 360 FOV lidars. This paper proposes an out-of-the-box NDT (Norma…

Cited by 1SourceScholar
2024

Orientation-Aware Multi-Modal Learning for Road Intersection Identification and Mapping

ICRA 2024poster

Accurate identification of road intersections is the pivotal task for automatic construction of high-definition maps, particularly in unstructured scenes. Existing methods predominantly rely on single-modal data and thus show an obvious unimodal limitation, i.e., lack of contextual information. More…

Cited by 2SourceScholar