← Search

Xingxing Zuo

29 accepted papers

2026

Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using A Feed-Forward 3D Model

RSS 2026poster

Fast and reliable initialization is critical for monocular visual–inertial navigation systems (VINS), as it establishes the starting conditions for subsequent state estimation. Despite steady progress, most existing methods heavily rely on visual feature correspondences and require 3-4 seconds of se…

Cited by 0SourceScholar
2026

EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation

CVPR 2026

Understanding and predicting object motion from egocentric video is fundamental to embodied perception and interaction. However, generating physically consistent 6DoF trajectories remains challenging due to occlusions, fast motion, and the lack of explicit physical reasoning in existing generative m

Cited by 0SourcecodeScholar
2026

MonoTher-Depth: Enhancing Thermal Depth Estimation Via Confidence-Aware Distillation

ICRA 2026poster

Monocular depth estimation (MDE) from thermal images is a crucial technology for robotic systems operating in challenging conditions such as fog, smoke, and low light. The limited availability of labeled thermal data constrains the generalization capabilities of thermal MDE models compared to founda…

2025

Gaussian-LIC: Real-Time Photo-Realistic SLAM with Gaussian Splatting and LiDAR-Inertial-Camera Fusion

ICRA 2025

In this paper, we present a real-time photo-realistic SLAM method based on marrying Gaussian Splatting with LiDAR-Inertial-Camera SLAM. Most existing radiance-field-based SLAM systems mainly focus on bounded indoor environments, equipped with RGB-D or RGB sensors. However, they are prone to decline

Cited by 28SourcecodeScholar
2025

L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR Model

IROS 2025

Semantic Scene Completion (SSC) constitutes a pivotal element in autonomous driving perception systems, tasked with inferring the 3D semantic occupancy of a scene from sensory data. To improve accuracy, prior research has implemented various computationally demanding and memory-intensive 3D operatio

Cited by 3SourcecodeScholar
2025

LiCROcc: Teach Radar for Accurate Semantic Occupancy Prediction Using LiDAR and Camera

RA-L 2025

Semantic Scene Completion (SSC) is pivotal in autonomous driving perception, frequently confronted with the complexities of weather and illumination changes. The long-term strategy involves fusing multi-modal information to bolster the system's robustness. Radar, increasingly utilized for 3D target

Cited by 18SourceScholar
2025

MonoTher-Depth: Enhancing Thermal Depth Estimation via Confidence-Aware Distillation

RA-L 2025

Monocular depth estimation (MDE) from thermal images is a crucial technology for robotic systems operating in challenging conditions such as fog, smoke, and low light. The limited availability of labeled thermal data constrains the generalization capabilities of thermal MDE models compared to founda

Cited by 2SourceScholar
2024

A Multimodal, Multi-Task Adapting Framework for Video Action Recognition

AAAI 2024technical

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance a…

Cited by 17SourcePDFScholar
2024

Caltech Aerial RGB-Thermal Dataset in the Wild

ECCV 2024poster

"We present the first publicly-available RGB-thermal dataset designed for aerial robotics operating in natural environments. Our dataset captures a variety of terrain across the United States, including rivers, lakes, coastlines, deserts, and forests, and consists of synchronized RGB, thermal, globa…

2024

Dynamic LiDAR Re-simulation using Compositional Neural Fields

CVPR 2024highlight

We introduce DyNFL a novel neural field-based approach for high-fidelity re-simulation of LiDAR scans in dynamic driving scenes. DyNFL processes LiDAR measurements from dynamic environments accompanied by bounding boxes of moving objects to construct an editable neural field. This field comprising s…

2024

LaPose: Laplacian Mixture Shape Modeling for RGB-Based Category-Level Object Pose Estimation

ECCV 2024poster

"While RGBD-based methods for category-level object pose estimation hold promise, their reliance on depth data limits their applicability in diverse scenarios. In response, recent efforts have turned to RGB-based methods; however, they face significant challenges stemming from the absence of depth i…

2024

RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale

ICRA 2024poster

We present a novel approach for metric dense depth estimation based on the fusion of a single-view image and a sparse, noisy Radar point cloud. The direct fusion of heterogeneous Radar and image data, or their encodings, tends to yield dense depth maps with significant artifacts, blurred boundaries,…

Cited by 11SourcecodeScholar
2023

Coco-LIC: Continuous-Time Tightly-Coupled LiDAR-Inertial-Camera Odometry Using Non-Uniform B-Spline

RA-L 2023

In this letter, we propose an efficient continuous-time LiDAR-Inertial-Camera Odometry, utilizing non-uniform B-splines to tightly couple measurements from the LiDAR, IMU, and camera. In contrast to uniform B-spline-based continuous-time methods, our non-uniform B-spline approach offers significant

Cited by 39SourcecodeScholar
2023

Incremental Dense Reconstruction From Monocular Video With Guided Sparse Feature Volume Fusion

RA-L 2023

Incrementally recovering 3D dense structures from monocular videos is of paramount importance since it enables various robotics and AR applications. Feature volumes have recently been shown to enable efficient and accurate incremental dense reconstruction without the need to first estimate depth, bu

Cited by 12SourceScholar
2022

Ctrl-VIO: Continuous-Time Visual-Inertial Odometry for Rolling Shutter Cameras

RA-L 2022

In this letter, we propose a probabilistic continuous-time visual-inertial odometry (VIO) for rolling shutter cameras. The continuous-time trajectory formulation naturally facilitates the fusion of asynchronized high-frequency IMU data and motion-distorted rolling shutter images. To prevent intracta

Cited by 27SourcecodeScholar
2022

Symmetry and Uncertainty-Aware Object SLAM for 6DoF Object Pose Estimation

CVPR 2022poster

We propose a keypoint-based object-level SLAM framework that can provide globally consistent 6DoF pose estimates for symmetric and asymmetric objects alike. To the best of our knowledge, our system is among the first to utilize the camera pose information from SLAM to provide prior knowledge for tra…

Cited by 51PDFcodeScholar
2022

Visual-Inertial SLAM with Tightly-Coupled Dropout-Tolerant GPS Fusion

IROS 2022poster

Robotic applications are continuously striving towards higher levels of autonomy. To achieve that goal, a highly robust and accurate state estimation is indispensable. Combining visual and inertial sensor modalities has proven to yield accurate and locally consistent results in short-term applicatio…

Cited by 20SourceScholar
2021

CLINS: Continuous-Time Trajectory Estimation for LiDAR-Inertial System

IROS 2021poster

In this paper, we propose a highly accurate continuous-time trajectory estimation framework dedicated to SLAM (Simultaneous Localization and Mapping) applications, which enables fuse high-frequency and asynchronous sensor data effectively. We apply the proposed framework in a 3D LiDAR-inertial syste…

Cited by 50SourcecodeScholar
2021

CodeVIO: Visual-Inertial Odometry with Learned Optimizable Dense Depth

ICRA 2021poster

In this work, we present a lightweight, tightly-coupled deep depth network and visual-inertial odometry (VIO) system, which can provide accurate state estimates and dense depth maps of the immediate surroundings. Leveraging the proposed lightweight Conditional Variational Autoencoder (CVAE) for dept…

Cited by 52SourceScholar
2020

LIC-Fusion 2.0: LiDAR-Inertial-Camera Odometry with Sliding-Window Plane-Feature Tracking

IROS 2020poster

Multi-sensor fusion of multi-modal measurements from commodity inertial, visual and LiDAR sensors to provide robust and accurate 6DOF pose estimation holds great potential in robotics and beyond. In this paper, building upon our prior work (i.e., LIC-Fusion), we develop a sliding-window filter based…

Cited by 150SourceScholar
2020

Targetless Calibration of LiDAR-IMU System Based on Continuous-time Batch Estimation

IROS 2020poster

Sensor calibration is the fundamental block for a multi-sensor fusion system. This paper presents an accurate and repeatable LiDAR-IMU calibration method (termed LI-Calib), to calibrate the 6-DOF extrinsic transformation between the 3D LiDAR and the Inertial Measurement Unit (IMU). Regarding the hig…

Cited by 94SourcecodeScholar
2019

Tightly-Coupled Aided Inertial Navigation with Point and Plane Features

ICRA 2019poster

This paper presents a tightly-coupled aided inertial navigation system (INS) with point and plane features, a general sensor fusion framework applicable to any visual and depth sensor (e.g., RGBD, LiDAR) configuration, in which the camera is used for point feature tracking and depth sensor for plane…

Cited by 53SourceScholar
2019

Visual-Inertial Localization With Prior LiDAR Map Constraints

RA-L 2019

In this letter, we develop a low-cost stereo visual-inertial localization system, which leverages efficient multi-state constraint Kalman filter (MSCKF)-based visual-inertial odometry (VIO) while utilizing an <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/

Cited by 59SourceScholar