← Search

Shenghai Yuan

69 accepted papers

2026

Accurate Calibration and Robust LiDAR-Inertial Odometry for Spinning Actuated LiDAR Systems

RA-L 2026

Accurate calibration and robust localization are fundamental for downstream tasks in spinning actuated LiDAR applications. Existing methods, however, require parameterizing extrinsic parameters based on different mounting configurations, limiting their generalizability. Additionally, spinning actuat

Cited by 0SourcecodeScholar
2026

Adaptive 3D Perception for Small Aerial Targets Under Sparse Sampling via Reinforcement Learning

CVPR 2026

Detecting small aerial targets (SATs) in long-range LiDAR is challenging because motion causes extreme variations in point density, breaking fixed-voxel and static-threshold assumptions in standard 3D detection and tracking. To address the challenges, we introduce A3PRL, an RL-driven adaptive percep

Cited by 0SourceScholar
2026

CETUS: Causal Event-Driven Temporal Modeling with Unified Variable-Rate Scheduling

ICRA 2026poster

Event cameras capture asynchronous pixel-level brightness changes with microsecond temporal resolution, offering unique advantages for high-speed vision tasks. Existing methods often convert event streams into intermediate representations such as frames, voxel grids, or point clouds, which inevitabl…

2026

ColorMap-VIO: A Drift-Free Visual-Inertial Odometry in a Prior Colored Point Cloud Map

RA-L 2026

Visual-inertial odometry (VIO) can estimate robot poses at high frequencies but suffers from accumulated drift over time. Incorporating point cloud maps offers a promising solution, yet existing registration methods between vision and point clouds are limited by heterogeneous feature alignment, leav

Cited by 0SourceScholar
2026

Comp-Attn: Present-and-Align Attention for Compositional Video Genneration

ICML 2026poster

In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main challenges are twofold: 1) Subject presence, where not all subjects can be presented in the video; 2) Inter-subject relat…

Cited by 0SourceScholar
2026

Following Is All You Need: Robot Crowd Navigation Using People As Planners

ICRA 2026poster

Navigating in crowded environments requires the robot to be equipped with high-level reasoning and planning techniques. Existing works focus on developing complex and heavyweight planners while ignoring the role of human intelligence. Since humans are highly capable agents who are also widely availa…

2026

Global Planning for Object Navigation Via a Weighted Traveling Repairman Problem Formulation

ICRA 2026poster

Zero-Shot Object Navigation (ZSON) requires agents to navigate to objects specified via open-ended natural language without predefined categories or prior environmental knowledge. While recent methods leverage foundation models or multi-modal maps, they often rely on 2D representations and greedy st…

Cited by 0codeScholar
2026

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

ICLR 2026poster

We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference s…

Cited by 0SourcecodeScholar
2026

PERAL: Perception-Aware Motion Control for Passive LiDAR Excitation in Spherical Robots

ICRA 2026poster

Autonomous mobile robots increasingly rely on LiDAR–IMU odometry for navigation and mapping, yet horizontally mounted LiDARs (e.g., MID360) capture limited near-ground returns, reducing terrain awareness and degrading performance in feature-scarce environments. Prior solutions, such as static tilt, …

2026

Proactive Risk-Aware Trajectory Planning for Autonomous Driving in Unstructured Environments Via Reinforcement Learning with Adaptive Reward Design

ICRA 2026poster

Trajectory planning for autonomous driving in dynamic unstructured traffic remains a fundamental challenge. Existing methods are often reactive, i.e., they only respond to observed situations without explicitly anticipating future risks. Moreover, most reinforcement learning based approaches rely on…

Cited by 0Scholar
2026

SplatSSC: Decoupled Depth-Guided Gaussian Splatting for Semantic Scene Completion

AAAI 2026technical

Monocular 3D Semantic Scene Completion (SSC) is a challenging yet promising task that aims to infer dense geometric and semantic descriptions of a scene from a single image. While recent object-centric paradigms significantly improve efficiency by leveraging flexible 3D Gaussian primitives, they sti

Cited by 0SourcePDFScholar
2026

TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion

CVPR 2026

Embodied 3D Semantic Scene Completion (SSC) infers dense geometry and semantics from continuous egocentric observations. Most existing Gaussian-based methods rely on random initialization of many primitives within predefined spatial bounds, resulting in redundancy and poor scalability to unbounded s

Cited by 0SourcecodeScholar
2025

AirSwarm: Enabling Cost-Effective Multi-UAV Research with COTS drones

IROS 2025

Traditional unmanned aerial vehicle (UAV) swarm missions rely heavily on expensive custom-made drones with onboard perception or external positioning systems, limiting their widespread adoption in research and education. To address this issue, we propose AirSwarm. AirSwarm democratizes multi-drone c

Cited by 4SourcecodeScholar
2025

Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling

ICASSP 2025accepted

As small unmanned aerial vehicles (UAVs) become increasingly prevalent, there is growing concern regarding their impact on public safety and privacy, highlighting the need for advanced tracking and trajectory estimation solutions. In response, this paper introduces a novel framework that utilizes au…

Cited by 0SourceScholar
2025

Autonomous 3D Moving Target Encirclement and Interception with Range Measurement

IROS 2025

Commercial UAVs are an emerging security threat as they are capable of carrying hazardous payloads or disrupting air traffic. To counter UAVs, we introduce an autonomous 3D target encirclement and interception strategy. Unlike traditional ground-guided systems, this strategy employs autonomous drone

Cited by 3SourceScholar
2025

BEV-LIO(LC): BEV Image Assisted LiDAR-Inertial Odometry with Loop Closure

IROS 2025

This work introduces BEV-LIO(LC), a novel LiDAR-Inertial Odometry (LIO) framework that combines Bird’s Eye View (BEV) image representations of LiDAR data with geometry-based point cloud registration and incorporates loop closure (LC) through BEV image features. By normalizing point density, we proje

Cited by 11SourcecodeScholar
2025

CGS-SLAM: Compact 3D Gaussian Splatting for Dense Visual SLAM

IROS 2025

Recent work has shown that 3D Gaussian-based SLAM enables high-quality reconstruction, accurate pose estimation, and real-time rendering of scenes. However, these approaches are built on a tremendous number of redundant 3D Gaussian ellipsoids, leading to high memory and storage costs and slow traini

Cited by 61SourceScholar
2025

DPGP: A Hybrid 2D-3D Dual Path Potential Ghost Probe Zone Prediction Framework for Safe Autonomous Driving

IROS 2025

Modern robots must coexist with humans in dense urban environments. A key challenge is the ghost probe problem, where pedestrians or objects unexpectedly rush into traffic paths. This issue affects both autonomous vehicles and human drivers. Existing works propose vehicle-to-everything (V2X) strateg

Cited by 2SourceScholar
2025

EGS-SLAM: RGB-D Gaussian Splatting SLAM With Events

RA-L 2025

Gaussian Splatting SLAM (GS-SLAM) offers a notable improvement over traditional SLAM methods, in enabling photorealistic 3D reconstruction that conventional approaches often struggle to achieve. However, existing GS-SLAM systems perform poorly under persistent and severe motion blur commonly encount

Cited by 3SourceScholar
2025

Enhancing Scene Coordinate Regression With Efficient Keypoint Detection and Sequential Information

RA-L 2025

Scene Coordinate Regression (SCR) is a visual localization technique that utilizes deep neural networks (DNN) to directly regress 2D-3D correspondences for camera pose estimation. However, current SCR methods often face challenges in handling repetitive textures and meaningless areas due to their re

Cited by 3SourcecodeScholar
2025

Following is All You Need: Robot Crowd Navigation Using People as Planners

RA-L 2025

Navigating in crowded environments requires the robot to be equipped with high-level reasoning and planning techniques. Existing works focus on developing complex and heavyweight planners while ignoring the role of human intelligence. Since humans are highly capable agents who are also widely availa

Cited by 3SourceScholar
2025

GERA: Geometric Embedding for Efficient Point Registration Analysis

ICRA 2025

Point cloud registration aims to provide estimated transformations to align point clouds, which plays a crucial role in pose estimation of various navigation systems, such as surgical guidance systems and autonomous vehicles. Despite the impressive performance of recent models on benchmark datasets,

Cited by 3SourceScholar
2025

HelmetPoser: A Helmet-Mounted IMU Dataset for Data-Driven Estimation of Human Head Motion in Diverse Conditions

ICRA 2025

Helmet-mounted wearable positioning systems are crucial for enhancing safety and facilitating coordination in industrial, construction, and emergency rescue environments. These systems, including LiDAR-Inertial Odometry (LIO) and Visual-Inertial Odometry (VIO), often face challenges in localization

Cited by 9SourcecodeScholar
2025

Identity-Preserving Text-to-Video Generation by Frequency Decomposition

CVPR 2025highlight

Identity-preserving text-to-video (IPT2V) generation aims to create high-fidelity videos with consistent human identity. It is an important task in video generation but remains an open problem for generative models. This paper pushes the technical frontier of IPT2V in two directions that have not b…

2025

ImgEdit: A Unified Image Editing Dataset and Benchmark

NeurIPS 2025poster

Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and insufficient benchmarks. To overcome these limitations, we introduce…

Cited by 0SourcecodeScholar
2025

Large-Scale UWB Anchor Calibration and One-Shot Localization Using Gaussian Process

ICRA 2025

Ultra-wideband (UWB) is gaining popularity with devices like AirTags for precise home item localization but faces significant challenges when scaled to large environments like seaports. The main challenges are calibration and localization under obstructed conditions, which are common in logistics en

Cited by 15SourceScholar
2025

LiMo-Calib: On-Site Fast LiDAR-Motor Calibration for Quadruped Robot-Based Panoramic 3D Sensing System

IROS 2025

Conventional single LiDAR systems are inherently constrained by their limited field of view (FoV), leading to blind spots and incomplete environmental awareness, particularly on robotic platforms with strict payload limitations. Integrating a motorized LiDAR offers a practical solution by significan

Cited by 11SourcecodeScholar
2025

MNE-SLAM: Multi-Agent Neural SLAM for Mobile Robots

CVPR 2025poster

Neural implicit scene representations have recently shown promising results in dense visual SLAM. However, existing implicit SLAM algorithms are constrained to single-agent scenarios, and fall difficulty in large indoor scenes and long sequences. Existing multi-agent SLAM frameworks cannot meet the…

2025

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

NeurIPS 2025poster

Subject-to-Video (S2V) generation aims to create videos that faithfully incorporate reference content, providing enhanced flexibility in the production of videos. To establish the infrastructure for S2V generation, we propose **OpenS2V-Nexus**, consisting of (i) **OpenS2V‑Eval**, a fine‑grained benc…

Cited by 0SourceScholar
2025

Realm: Real-Time Line-of-Sight Maintenance in Multi-Robot Navigation with Unknown Obstacles

ICRA 2025

Multi-robot navigation in complex environments relies on inter-robot communication and mutual observation for situational awareness. This paper studies the multi-robot navigation problem in unknown environments with line-ofsight (LoS) connectivity constraints. While previous works are limited to kno

Cited by 8SourcecodeScholar
2025

Robust Loop Closure by Textual Cues in Challenging Environments

RA-L 2025

Loop closure is an important task in robot navigation. However, existing methods mostly rely on some implicit or heuristic features of the environment, which can still fail to work in common environments such as corridors, tunnels, and warehouses. Indeed, navigating in such featureless, degenerative

Cited by 12SourcecodeScholar
2025

Swept Volume-Aware Trajectory Planning and MPC Tracking for Multi-Axle Swerve-Drive AMRs

ICRA 2025

Multi-axle autonomous mobile robots (AMRs) are set to revolutionize the future of robotics in logistics. As the backbone of next-generation solutions, these robots face a critical challenge: managing and minimizing swept volume during turns while maintaining precise control. Traditional systems desi

Cited by 5SourceScholar
2025

UA-MPC: Uncertainty-Aware Model Predictive Control for Motorized LiDAR Odometry

RA-L 2025

Accurate and comprehensive 3D sensing using LiDAR systems is crucial for various applications in photogrammetry and robotics, including facility inspection, Building Information Modeling (BIM), and robot navigation. Motorized LiDAR systems can expand the Field of View (FoV) without adding multiple s

Cited by 33SourcecodeScholar
2025

UAVScenes: A Multi-Modal Dataset for UAVs

ICCV 2025poster

Multi-modal perception is essential for unmanned aerial vehicle (UAV) operations, as it enables a comprehensive understanding of the UAVs' surrounding environment. However, most existing multi-modal UAV datasets are primarily biased toward localization and 3D reconstruction tasks, or only support ma…

2025

ULOC: Learning to Localize in Complex Large-Scale Environments with Ultra-Wideband Ranges

ICRA 2025

While UWB-based methods can achieve high localization accuracy in small-scale areas, their accuracy and reliability are significantly challenged in large-scale environments. In this paper, we propose a learning-based framework named ULOC for Ultra-Wideband (UWB) based localization in such complex, l

Cited by 11SourcecodeScholar
2025

Underwater target 6D State Estimation via UUV Attitude Enhance Observability

IROS 2025

Accurate relative state observation of Unmanned Underwater Vehicles (UUVs) for tracking uncooperative targets remains a significant challenge due to the absence of GPS, complex underwater dynamics, and sensor limitations. Existing localization approaches rely on either global positioning infrastruct

Cited by 1SourceScholar
2025

Unsupervised UAV 3D Trajectories Estimation with Sparse Point Clouds

ICASSP 2025accepted

Compact UAV systems, while advancing delivery and surveillance, pose significant security challenges due to their small size, which hinders detection by traditional methods. This paper presents a cost-effective, unsupervised UAV detection method using spatial-temporal sequence processing to fuse mul…

Cited by 0SourceScholar
2025

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

CVPR 2025poster

Video Variational Autoencoder (VAE) encodes videos into a low-dimensional latent space, becoming a key component of most Latent Video Diffusion Models (LVDMs) to reduce model training costs. However, as the resolution and duration of generated videos increase, the encoding cost of Video VAEs becomes…

2024

ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation

NeurIPS 2024spotlight

We propose a novel text-to-video (T2V) generation benchmark, *ChronoMagic-Bench*, to evaluate the temporal and metamorphic knowledge skills in time-lapse video generation of the T2V models (e.g. Sora and Lumiere). Compared to existing benchmarks that focus on visual quality and text relevance of gen…

2024

Eigen Is All You Need: Efficient Lidar-Inertial Continuous-Time Odometry With Internal Association

RA-L 2024

In this paper, we propose a continuous-time lidar-inertial odometry (CT-LIO) system named SLICT2, which promotes two main insights. One, contrary to conventional wisdom, CT-LIO algorithm can be optimized by linear solvers in only a few iterations, which is more efficient than commonly used nonlinear

Cited by 23SourcecodeScholar
2024

I2EKF-LO: A Dual-Iteration Extended Kalman Filter Based LiDAR Odometry

IROS 2024poster

LiDAR odometry is a pivotal technology in the fields of autonomous driving and autonomous mobile robotics. However, most of the current works focus on nonlinear optimization methods, and still existing many challenges in using the traditional Iterative Extended Kalman Filter (IEKF) framework to tack…

Cited by 10SourcecodeScholar
2024

LIO-GVM: An Accurate, Tightly-Coupled Lidar-Inertial Odometry With Gaussian Voxel Map

RA-L 2024

This letter presents a probabilistic voxel-based LiDAR Inertial Odometry framework for accurate and robust pose estimation. The framework addresses the correspondence mismatching issue by representing the LiDAR points as a set of Gaussian distributions and evaluating the divergence in variance for o

Cited by 21SourcecodeScholar
2024

MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception

CVPR 2024highlight

Perception plays a crucial role in various robot applications. However existing well-annotated datasets are biased towards autonomous driving scenarios while unlabelled SLAM datasets are quickly over-fitted and often lack environment and domain variations. To expand the frontier of these fields we i…

Cited by 36SourcePDFScholar
2024

MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone Threats

ICRA 2024poster

In response to the evolving challenges posed by small unmanned aerial vehicles (UAVs), which possess the potential to transport harmful payloads or independently cause damage, we introduce MMAUD: a comprehensive Multi-Modal Anti-UAV Dataset. MMAUD addresses a critical gap in contemporary threat dete…

Cited by 22SourcecodeScholar
2024

MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation

ICRA 2024poster

Multi-modal unsupervised domain adaptation (MM-UDA) for 3D semantic segmentation is a practical solution to embed semantic understanding in autonomous systems without expensive point-wise annotations. While previous MM-UDA methods can achieve overall improvement, they suffer from significant class-i…

Cited by 19SourcecodeScholar
2024

Multi-Robot Active Graph Exploration with Reduced Pose-SLAM Uncertainty via Submodular Optimization

IROS 2024poster

This paper considers the multi-robot active graph exploration problem, where robots need to collaboratively cover a graph environment while maintaining reliable pose estimation in collaborative Simultaneous Localization and Mapping (SLAM). Considering both objectives presents challenges for multi-ro…

Cited by 2SourcecodeScholar
2024

Outram: One-shot Global Localization via Triangulated Scene Graph and Global Outlier Pruning

ICRA 2024poster

One-shot LiDAR localization refers to the ability to estimate the robot pose from one single point cloud, which yields significant advantages in initialization and relocalization processes. In the point cloud domain, the topic has been extensively studied as a global descriptor retrieval (i.e., loop…

Cited by 19SourcecodeScholar
2024

PSS-BA: LiDAR Bundle Adjustment with Progressive Spatial Smoothing

IROS 2024poster

Accurate and consistent construction of point clouds from LiDAR scanning data is fundamental for 3D modeling applications. Current solutions, such as multiview point cloud registration and LiDAR bundle adjustment, predominantly depend on the local plane assumption, which may be inadequate in complex…

Cited by 10SourceScholar
2024

Reliable Spatial-Temporal Voxels For Multi-Modal Test-Time Adaptation

ECCV 2024poster

"Multi-modal test-time adaptation (MM-TTA) is proposed to adapt models to an unlabeled target domain by leveraging the complementary multi-modal inputs in an online manner. Previous MM-TTA methods for 3D segmentation rely on predictions of cross-modal information in each input frame, while they igno…

2024

SGBA: Semantic Gaussian Mixture Model-Based LiDAR Bundle Adjustment

RA-L 2024

LiDAR bundle adjustment (BA) is an effective approach to reduce the drifts in pose estimation from the front-end. Existing works on LiDAR BA usually rely on predefined geometric features for landmark representation. This reliance restricts generalizability, as the system will inevitably deteriorate

Cited by 8SourceScholar
2024

Salient Sparse Visual Odometry With Pose-Only Supervision

RA-L 2024

Visual Odometry (VO) is vital for the navigation of autonomous systems, providing accurate position and orientation estimates at reasonable costs. While traditional VO methods excel in some conditions, they struggle with challenges like variable lighting and motion blur. Deep learning-based VO, thou

Cited by 14SourceScholar
2023

AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian Awareness

IROS 2023poster

In this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications. However, traditional approaches that rely on cameras and LID…

Cited by 9SourcecodeScholar
2023

DoubleBee: A Hybrid Aerial-Ground Robot with Two Active Wheels

IROS 2023poster

In this paper, we present the dynamic model and control of DoubleBee, a novel hybrid aerial-ground vehicle consisting of two propellers mounted on tilting servo motors and two motor-driven wheels. DoubleBee exploits the high energy efficiency of a bicopter configuration in aerial mode, and enjoys th…

Cited by 19SourceScholar
2023

MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensing

NeurIPS 2023poster

4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensi…

2023

Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation

ICCV 2023poster

Continual Test-Time Adaptation (CTTA) generalizes conventional Test-Time Adaptation (TTA) by assuming that the target domain is dynamic over time rather than stationary. In this paper, we explore Multi-Modal Continual Test-Time Adaptation (MM-CTTA) as a new extension of CTTA for 3D semantic segmenta…

Cited by 20PDFScholar
2023

Non-cooperative Stochastic Target Encirclement by Anti-synchronization Control via Range-only Measurement

ICRA 2023poster

This paper investigates the stochastic moving target encirclement problem in a realistic setting. In contrast to typical assumptions in related works, the target in our work is non-cooperative and capable of escaping the circle containment by boosting its speed to maximum for a short duration. In ex…

Cited by 12SourceScholar
2023

Path Planning for Multiple Tethered Robots Using Topological Braids

RSS 2023poster

Path planning for multiple tethered robots is a challenging problem due to the complex interactions among the cables and the possibility of severe entanglements. Previous works on this problem either consider idealistic cable models or provide no guarantee for entanglement-free paths. In this work,…

2023

SLICT: Multi-Input Multi-Scale Surfel-Based Lidar-Inertial Continuous-Time Odometry and Mapping

RA-L 2023

While feature association to a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">global map</i> has significant benefits, to keep the computations from growing exponentially, most lidar-based odometry and mapping methods opt to associate features with

Cited by 58SourcecodeScholar
2023

Segregator: Global Point Cloud Registration with Semantic and Geometric Cues

ICRA 2023poster

This paper presents Segregator, a global point cloud registration framework that exploits both semantic information and geometric distribution to efficiently build up outlier-robust correspondences and search for inliers. Current state-of-the-art algorithms rely on point features to set up putative…

Cited by 28SourcecodeScholar
2022

DIRECT: A Differential Dynamic Programming Based Framework for Trajectory Generation

RA-L 2022

This letter introduces a differential dynamic programming (DDP) based framework for polynomial trajectory generation for differentially flat systems. In particular, instead of using a linear equation with increasing size to represent multiple polynomial segments as in literature, we take a new persp

Cited by 20SourcecodeScholar
2021

LIRO: Tightly Coupled Lidar-Inertia-Ranging Odometry

ICRA 2021poster

In recent years, thanks to the continuously reduced cost and weight of 3D lidar, the applications of this type of sensor in the community have become increasingly popular. Despite many progresses, estimation drift and tracking loss are still prevalent concerns associated with these systems. However,…

Cited by 41SourceScholar
2021

MILIOM: Tightly Coupled Multi-Input Lidar-Inertia Odometry and Mapping

RA-L 2021

In this letter we investigate a tightly coupled Lidar-Inertia Odometry and Mapping (LIOM) scheme, with the capability to incorporate multiple lidars with complementary field of view (FOV). In essence, we devise a time-synchronized scheme to combine extracted features from separate lidars into a sing

Cited by 45SourceScholar
2020

OriNet: Robust 3-D Orientation Estimation With a Single Particular IMU

RA-L 2020

Estimating the robot's heading is a crucial requirement in odometry systems which are attempting to estimate the movement trajectory of a robot. Small errors in the orientation estimation result in a significant difference between the estimated and real trajectory, and failure of the odometry system

Cited by 94SourceScholar