← Search

Xianghui Pan

6 accepted papers

2026

BEVDrive-E2E: Imitation With Bird's Eye View Perception for Interpretable End-to-End Autonomous Driving

RA-L 2026

Imitation learning (IL) for end-to-end autonomous driving (E2E-AD) has made great progress recently in the closed-loop evaluation of the CARLA simulator. However, the causal confusion remains an open problem. To address this issue, we propose the BEVDrive-E2E to explore the interpretability of the e

Cited by 0SourceScholar
2026

Dynamics Are Learned, Not Told: Semi-Supervised Discovery of Latent Dynamics Geometries For Zero-Shot Policy Adaptation

ICML 2026poster

Real-world dynamics shifts pose a critical challenge for reinforcement learning, yet prior methods typically rely on encoding explicitly identified physical parameters into a latent context, a rigid parameterization that proves brittle to unmodeled or compound dynamics variations. We instead investi…

Cited by 0SourceScholar
2025

LGPR: Local Feature Learning Brings More Generalizable Visual Place Recognition

IROS 2025

We propose a Visual Place Recognition (VPR) framework by sharing lightweight keypoint extraction modules for local features. Current research on the joint learning of local keypoint matching and VPR is relatively scarce, and the application deployment of real-time spatial computing on edge devices h

Cited by 0SourcecodeScholar
2025

Rotation-Equivariant Robot Vision: A Perspective via Correspondence-Matching and Pre-training

IROS 2025

Correspondence matching is a fundamental and crucial task in robot vision. In recent years, deep learning-based keypoint matching techniques have shown outstanding performance in downstream tasks. Conventional learning-based correspondence matching methods rely on large datasets and a specific train

Cited by 0SourceScholar
2024

DVT: Decoupled Dual-Branch View Transformation for Monocular Bird’s Eye View Semantic Segmentation

IROS 2024poster

Monocular Bird’s Eye View (BEV) semantic segmentation is critical for autonomous driving for its inherent advantages in spatial representation and downstream tasks. However, it is challenging to simultaneously learn view transformation and pixel-wise classification. Previous works suffer from non-fl…

Cited by 0SourcecodeScholar
2024

GenerOcc: Self-supervised Framework of Real-time 3D Occupancy Prediction for Monocular Generic Cameras

IROS 2024poster

In the context of 3D scene perception tasks, the significance of 3D occupancy prediction has been progressively growing, aiming to forecast the occupancy state of voxels in a discrete 3D space. However, existing methods typically exhibit several limitations, such as restricted adaptability to non-pi…

Cited by 0SourceScholar