← Search

Jialei Xu

3 accepted papers

2026

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models

CVPR 2026

Large vision-language models (VLMs) show strong multimodal understanding but still struggle with 3D spatial reasoning, such as distance estimation, size comparison, and cross-view consistency. Existing 3D-aware methods either depend on auxiliary 3D information or enhance RGB-only VLMs with geometry

Cited by 0SourceScholar
2025

Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios

ICRA 2025

Monocular depth estimation from RGB images plays a pivotal role in 3D vision. However, its accuracy can deteriorate in challenging environments such as nighttime or adverse weather conditions. While long-wave infrared cameras offer stable imaging in such challenging conditions, they are inherently l

Cited by 6SourceScholar
2024

SDGE: Stereo Guided Depth Estimation for 360°Camera Sets

IROS 2024poster

Depth estimation is a critical technology in autonomous driving, and multi-camera systems are often used to achieve a 360° perception. These 360° camera sets often have limited or low-quality overlap regions, making multi-view stereo methods infeasible for the entire image. Alternatively, monocular…

Cited by 1SourcecodeScholar