← Search

Xiaokai Bai

7 accepted papers

2026

Learning Better UAV-Based Cross-View Object Geo-Localization from Multi-Modal Prompts: MoP-UAV Benchmark and MoPT Framework

AAAI 2026technical

We present MoP-UAV, a new benchmark for UAV-based cross-view object geo-localization guided by multi-modal prompts. MoP-UAV supports fine-grained object-level cross-view localization under diverse prompt modalities, including natural language, bounding boxes, and click points. It offers potential fo

Cited by 0SourcePDFScholar
2026

RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection

CVPR 2026

4D millimeter-wave radar is a promising sensing modality for autonomous driving, yet effective 3D object detection from 4D radar and monocular images remains challenging. Existing fusion approaches either rely on instance proposals lacking global context or dense BEV grids constrained by rigid struc

Cited by 0SourcecodeScholar
2025

Boosting Multi-View Indoor 3D Object Detection via Adaptive 3D Volume Construction

ICCV 2025poster

This work presents SGCDet, a novel multi-view indoor 3D object detection framework based on adaptive 3D volume construction. Unlike previous approaches that restrict the receptive field of voxels to fixed locations on images, we introduce a geometry and context aware aggregation module to integrate…

2025

LGDD: Local-Global Synergistic Dual-Branch 3D Object Detection Using 4D Radar

IROS 2025

4D millimeter-wave radar plays a pivotal role in autonomous driving due to its cost-effectiveness and robustness in adverse weather. However, the application of 4D radar point cloud in 3D perception tasks is hindered by its inherent sparsity and noise. To address these challenges, we propose LGDD, a

Cited by 3SourcecodeScholar
2025

S-BEVLoc: BEV-Based Self-Supervised Framework for Large-Scale LiDAR Global Localization

RA-L 2025

LiDAR-based global localization is an essential component of simultaneous localization and mapping (SLAM), which helps loop closure and re-localization. Current approaches rely on ground-truth poses obtained from GPS or SLAM odometry to supervise network training. Despite the great success of these

Cited by 0SourceScholar
2025

SGDet3D: Semantics and Geometry Fusion for 3D Object Detection Using 4D Radar and Camera

RA-L 2025

4D millimeter-wave radar has gained attention as an emerging sensor for autonomous driving in recent years. However, existing 4D radar and camera fusion models often fail to fully exploit complementary information within each modality and lack deep cross-modal interactions. To address these issues,

Cited by 26SourceScholar