← Search

Yixi Cai

17 accepted papers

2026

DoGFlow: Self-Supervised LiDAR Scene Flow via Cross-Modal Doppler Guidance

RA-L 2026

Accurate 3D scene flow estimation is critical for autonomous systems to navigate dynamic environments safely, but creating the necessary large-scale, manually annotated datasets remains a significant bottleneck for developing robust perception models. Current self-supervised methods struggle to matc

Cited by 3SourcecodeScholar
2026

Learning to Localize Reference Trajectories in Image-Space for Visual Navigation

RSS 2026poster

We present LoTIS, a model for visual navigation that provides robot-agnostic image-space guidance by localizing a reference RGB trajectory in the robot’s current view, without requiring camera calibration, poses, or robot-specific training. Instead of predicting actions tied to specific robots, we p…

Cited by 0SourceScholar
2026

PRIX: Learning to Plan From Raw Pixels for End-to-End Autonomous Driving

RA-L 2026

While end-to-end autonomous driving models show promising results, their practical deployment is often hindered by large model sizes, a reliance on expensive LiDAR sensors and computationally intensive BEV feature representations. This limits their scalability, especially for mass-market vehicles eq

Cited by 10SourceScholar
2026

ProbPer-LiLo: Probabilistic Persistency Modeling for Life-Long Mapping

ICRA 2026poster

3D mapping is vital for a broad range of applications that rely on a consistent and accurate representation of the environment. Change is an ever-persistent force in our world and with the evolution of a scene its 3D map becomes outdated. Thus, a mapping framework that can adapt and refine the 3D ma…

Cited by 0SourceScholar
2025

DeltaFlow: An Efficient Multi-frame Scene Flow Estimation Method

NeurIPS 2025spotlight

Previous dominant methods for scene flow estimation focus mainly on input from two consecutive frames, neglecting valuable information in the temporal domain. While recent trends shift towards multi-frame reasoning, they suffer from rapidly escalating computational costs as the number of frames grow…

Cited by 0SourcecodeScholar
2025

Efficient Swept Volume-Based Trajectory Generation for Arbitrary-Shaped Ground Robot Navigation

IROS 2025

Navigating an arbitrary-shaped ground robot safely in cluttered environments remains a challenging problem. The existing trajectory planners that account for the robot’s physical geometry severely suffer from the intractable runtime. To achieve both computational efficiency and Continuous Collision

Cited by 0SourceScholar
2025

FAST-LIVO2 on Resource-Constrained Platforms: LiDAR-Inertial-Visual Odometry With Efficient Memory and Computation

RA-L 2025

This paper presents a lightweight LiDAR-inertial-visual odometry system optimized for resource-constrained platforms. It integrates a degeneration-aware adaptive visual frame selector into error-state iterated Kalman filter (ESIKF) with sequential updates, improving computation efficiency markedly w

Cited by 5SourceScholar
2025

GauSS-MI: Gaussian Splatting Shannon Mutual Information for Active 3D Reconstruction

RSS 2025poster

This research tackles the challenge of real-time active view selection and uncertainty quantification on visual quality for active 3D reconstruction. Visual quality is a critical aspect of 3D reconstruction. Recent advancements such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) h…

Cited by 0PDFcodeScholar
2025

LVBA: LiDAR-Visual Bundle Adjustment for RGB Point Cloud Mapping

ICRA 2025

Point cloud maps with accurate color are crucial in robotics and mapping applications. Existing approaches for producing RGB-colorized maps are primarily based on realtime localization using filter-based estimation or sliding window optimization, which may lack accuracy and global consistency. In th

Cited by 2SourceScholar
2025

Neural Surface Reconstruction and Rendering for LiDAR-Visual Systems

ICRA 2025

This paper presents a unified surface reconstruction and rendering framework for LiDAR-visual systems, integrating Neural Radiance Fields (NeRF) and Neural Distance Fields (NDF) to recover both appearance and structural information from posed images and point clouds. We address the structural visibl

Cited by 5SourcecodeScholar
2025

Temporal Overlapping Prediction: A Self-supervised Pre-training Method for LiDAR Moving Object Segmentation

ICCV 2025poster

Moving object segmentation (MOS) on LiDAR point clouds is crucial for autonomous systems such as self-driving vehicles. While previous supervised approaches rely on costly manual annotations, LiDAR sequences naturally capture temporal motion cues that can be leveraged for self-supervised learning. I…

2024

ROG-Map: An Efficient Robocentric Occupancy Grid Map for Large-scene and High-resolution LiDAR-based Motion Planning

IROS 2024poster

Recent advances in LiDAR technology have opened up new possibilities for robotic navigation. Given the widespread use of occupancy grid maps (OGMs) in robotic motion planning, this paper aims to address the challenges of integrating LiDAR with OGMs. To this end, we propose ROG-Map, a uniform grid-ba…

Cited by 21SourcecodeScholar
2023

MARSIM: A Light-Weight Point-Realistic Simulator for LiDAR-Based UAVs

RA-L 2023

The emergence of low-cost, small form factor and light-weight solid-state LiDAR sensors have brought new opportunities for autonomous unmanned aerial vehicles (UAVs) by advancing navigation safety and computation efficiency. Yet the successful developments of LiDAR-based UAVs must rely on extensive

Cited by 56SourcecodeScholar