← Search

Wenchao Ding

36 accepted papers

2026

CMoE: Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots

ICRA 2026poster

For effective deployment in real-world environments, humanoid robots must autonomously navigate a diverse range of complex terrains with abrupt transitions. While the Vanilla mixture of experts (MoE) framework is theoretically capable of modeling diverse terrain features, in practice, the gating net…

2026

Collaborative Learning of Local 3D Occupancy Prediction and Versatile Global Occupancy Mapping

ICRA 2026poster

Vision-based 3D semantic occupancy prediction is vital for autonomous driving, enabling unified modeling of static infrastructure and dynamic agents. Global occupancy maps serve as long-term memory priors, providing valuable historical context that enhances local perception. This is particularly imp…

2026

DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulation

CVPR 2026

Robotic manipulation from visual observations remains challenging due to the lack of 3D consistent representations that can generalize across diverse viewpoints and sensor configurations. Existing approaches often rely on masked autoencoders or neural scene representations, which fail to capture cro

Cited by 0SourceScholar
2026

Drive in Corridors: Enhancing the Safety of End-To-End Autonomous Driving Via Corridor Learning and Planning

ICRA 2026poster

Safety remains one of the most critical challenges in autonomous driving systems. In recent years, the end-to-end driving has shown great promise in advancing vehicle autonomy in a scalable manner. However, existing approaches often face safety risks due to the lack of explicit behavior constraints.…

2026

DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios

ICRA 2026poster

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust pose estimation models. This is largely attributed to the inher…

2026

Flash-Mono: Feed-Forward Accelerated Gaussian Splatting Monocular SLAM

ICLR 2026poster

Monocular 3D Gaussian Splatting SLAM suffers from critical limitations in time efficiency, geometric accuracy, and multi-view consistency. These issues stem from the time-consuming $\textit{Train-from-Scratch}$ optimization and the lack of inter-frame scale consistency from single-frame geometry pri…

Cited by 0SourceScholar
2026

Learning High-Frequency Continuous Action Chunks in Latent Space

ICML 2026poster

Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal consistency at moderate action frequencies, it becomes insufficient when the action frequency is further increased (e.g., to 60~Hz). At such high frequ…

Cited by 0SourceScholar
2026

Learning a Unified Risk Map for Autonomous Driving in Partially Observable Environments

RA-L 2026

Occlusion-aware prediction remains a critical challenge in autonomous driving due to the inherent uncertainty of unobserved regions. Existing approaches either overestimate risk based on reachable states or struggle to predict accurate trajectories under high occlusion uncertainty. To address these

Cited by 0SourceScholar
2026

OccLLaMA: A Unified Occupancy-Language-Action World Model for Enhancing Motion Planning Via Multi-Task Learning

ICRA 2026poster

Scene understanding via multi-modal large language models and scene forecasting with world models have advanced the development of autonomous driving. The former maps visual inputs to driving-specific outputs, neglecting spatial reasoning and world dynamics. The latter captures world dynamics, lacki…

Cited by 0codeScholar
2026

Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids

RSS 2026poster

Realizing interactive whole-body control for multi-humanoid systems is critical for unlocking complex collaborative capabilities in shared environments. Although recent advancements have significantly enhanced the agility of individual robots, bridging the gap to physically coupled multi-humanoid in…

Cited by 1SourceScholar
2026

Seeing Is Believing: Grounding Long-Video Understanding in Spatio-Temporal Visual Evidence

AAAI 2026technical

Although Vision Language Models (VLMs) have excelled at image and video understanding, applying them to hour-long videos is held back by two interrelated challenges: exorbitant computational expense and a qualitative breakdown in long-term temporal reasoning. Thus, models tend to generate answers ba

Cited by 0SourcePDFScholar
2026

SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned Prediction

CVPR 2026

Recent progress in feed-forward 3D Gaussian Splatting (3DGS) has notably improved rendering quality. However, the spatially uniform and highly redundant 3DGS map generated by previous feed-forward 3DGS methods limits their integration into downstream reconstruction tasks. We propose SparseSplat, the

Cited by 0SourcecodeScholar
2026

Unveiling the Surprising Efficacy of Navigation Understanding in End-To-End Autonomous Driving

ICRA 2026poster

Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-end autonomous driving systems tend to over-rely on local scene understanding while failing to utilize global navigation i…

2026

VINGS-Mono: Visual-Inertial Gaussian Splatting Monocular SLAM in Large Scenes

ICRA 2026poster

VINGS-Mono is a monocular inertial Gaussian Splatting (GS) SLAM framework designed for large-scale scenes. It integrates four main components: VIO Front End, 2D Gaussian Map, NVS Loop Closure, and Dynamic Eraser. The VIO Front End processes RGB frames with dense bundle adjustment and uncertainty est…

2026

VistaBot: View-Robust Robot Manipulation Via Spatiotemporal-Aware View Synthesis

ICRA 2026poster

Recently, end-to-end robotic manipulation models have gained significant attention for their generalizability and scalability. However, they often suffer from limited robustness to camera viewpoint changes when training with a fixed camera. In this paper, we propose VistaBot, a novel framework that …

2025

Drive in Corridors: Enhancing the Safety of End-to-End Autonomous Driving via Corridor Learning and Planning

RA-L 2025

Safety remains one of the most critical challenges in autonomous driving systems. In recent years, the end-to-end driving has shown great promise in advancing vehicle autonomy in a scalable manner. However, existing approaches often face safety risks due to the lack of explicit behavior constraints.

Cited by 3SourcecodeScholar
2025

Dual-AEB: Synergizing Rule-Based and Multimodal Large Language Models for Effective Emergency Braking

ICRA 2025

Automatic Emergency Braking (AEB) systems are a crucial component in ensuring the safety of passengers in autonomous vehicles. Conventional AEB systems primarily rely on closed-set perception modules to recognize traffic conditions and assess collision risks. To enhance the adaptability of AEB syste

Cited by 3SourcecodeScholar
2025

DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios

RA-L 2025

In the realm of object pose estimation, scenarios involving both dynamic objects and moving cameras are prevalent. However, the scarcity of corresponding real-world datasets significantly hinders the development and evaluation of robust pose estimation models. This is largely attributed to the inher

Cited by 0SourceScholar
2025

Enhancing Indoor Occupancy Prediction via Sparse Query-Based Multi-Level Consistent Knowledge Distillation

RA-L 2025

Occupancy prediction provides critical geometric and semantic understanding for robotics but faces efficiency-accuracy trade-offs. Current dense methods suffer computational waste on empty voxels, while sparse query-based approaches lack robustness in diverse and complex indoor scenes. In this paper

Cited by 1SourceScholar
2025

HGS-Planner: Hierarchical Planning Framework for Active Scene Reconstruction Using 3D Gaussian Splatting

ICRA 2025

In complex missions such as search and rescue, robots must make intelligent decisions in unknown environments, relying on their ability to perceive and understand their surroundings. High-quality and real-time reconstruction enhances situational awareness and is crucial for intelligent robotics. Tra

Cited by 19SourceScholar
2025

Topology-Driven Trajectory Optimization for Modelling Controllable Interactions Within Multi-Vehicle Scenario

IROS 2025

Trajectory optimization in multi-vehicle scenarios faces challenges due to its non-linear, non-convex properties and sensitivity to initial values, making interactions between vehicles difficult to control. In this paper, inspired by topological planning, we propose a differentiable local homotopy i

Cited by 0SourceScholar
2024

DeepPointMap: Advancing LiDAR SLAM with Unified Neural Descriptors

AAAI 2024technical

Point clouds have shown significant potential in various domains, including Simultaneous Localization and Mapping (SLAM). However, existing approaches either rely on dense point clouds to achieve high localization accuracy or use generalized descriptors to reduce map size. Unfortunately, these two a…

2024

HGS-Mapping: Online Dense Mapping Using Hybrid Gaussian Representation in Urban Scenes

RA-L 2024

Online dense mapping of urban scenes forms a fundamental cornerstone for scene understanding and navigation of autonomous vehicles. Recent advancements in dense mapping methods are mainly based on NeRF, whose rendering speed is too slow to meet online requirements. 3D Gaussian Splatting (3DGS), with

Cited by 20SourceScholar
2024

OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data

ICRA 2024poster

In the era of big data and large models, automatic annotating functions for multi-modal data are of great significance for real-world AI-driven applications, such as autonomous driving and embodied AI. Unlike traditional closed-set annotation, open-vocabulary annotation is essential to achieve human…

Cited by 13SourcecodeScholar
2024

Swift-Mapping: Online Neural Implicit Dense Mapping in Urban Scenes

AAAI 2024technical

Online dense mapping of urban scenes is of paramount importance for scene understanding of autonomous navigation. Traditional online dense mapping methods fuse sensor measurements (vision, lidar, etc.) across time and space via explicit geometric correspondence. Recently, NeRF-based methods have pro…

Cited by 2SourcePDFScholar
2023

FlowMap: Path Generation for Automated Vehicles in Open Space Using Traffic Flow

ICRA 2023poster

There is extensive literature on perceiving road structures by fusing various sensor inputs such as lidar point clouds and camera images using deep neural nets. Leveraging the latest advance of neural architects (such as transformers) and bird-eye-view (BEV) representation, the road cognition accura…

Cited by 4SourceScholar
2023

Traffic Flow-Based Crowdsourced Mapping in Complex Urban Scenario

RA-L 2023

An accurate road topological structure is of great importance for autonomous driving in complex urban environments. Currently, most autonomous vehicles highly rely on the High-Definition map (HD map) to cruise across the city. Without the prior map, it's hard for vehicles to find right-turning and l

Cited by 13SourceScholar
2021

Learning to Predict Vehicle Trajectories with Model-based Planning

CoRL 2021poster

Predicting the future trajectories of on-road vehicles is critical for autonomous driving. In this paper, we introduce a novel prediction framework called PRIME, which stands for Prediction with Model-based Planning. Unlike recent prediction works that utilize neural networks to model scene context…

Cited by 160SourceScholar
2020

Efficient Uncertainty-aware Decision-making for Automated Driving Using Guided Branching

ICRA 2020poster

Decision-making in dense traffic scenarios is challenging for automated vehicles (AVs) due to potentially stochastic behaviors of other traffic participants and perception uncertainties (e.g., tracking noise and prediction errors, etc.). Although the partially observable Markov decision process (POM…

Cited by 60SourcecodeScholar
2020

PiP: Planning-informed Trajectory Prediction for Autonomous Driving

ECCV 2020poster

It is critical to predict the motion of surrounding vehicles for self-driving planning, especially in a socially compliant and flexible way. However, future prediction is challenging due to the interaction and uncertainty in driving behaviors. We propose planning-informed trajectory prediction (PiP)…

2019

Online Vehicle Trajectory Prediction using Policy Anticipation Network and optimization-based Context Reasoning

ICRA 2019poster

In this paper, we present an online two-level vehicle trajectory prediction framework for urban autonomous driving where there are complex contextual factors, such as lane geometries, road constructions, traffic regulations and moving agents. Our method combines high-level policy anticipation with l…

Cited by 84SourceScholar
2019

Predicting Vehicle Behaviors Over An Extended Horizon Using Behavior Interaction Network

ICRA 2019poster

Anticipating possible behaviors of traffic participants is an essential capability of autonomous vehicles. Many behavior detection and maneuver recognition methods only have a very limited prediction horizon that leaves inadequate time and space for planning. To avoid unsatisfactory reactive decisio…

Cited by 120SourceScholar
2019

Safe Trajectory Generation for Complex Urban Environments Using Spatio-Temporal Semantic Corridor

RA-L 2019

Planning safe trajectories for autonomous vehicles in complex urban environments is challenging since there are numerous semantic elements (such as dynamic agents, traffic lights, and speed limits) to consider. These semantic elements may have different mathematical descriptions, such as obstacle, c

Cited by 134SourcecodeScholar
2018

Trajectory Replanning for Quadrotors Using Kinodynamic Search and Elastic Optimization

ICRA 2018poster

We focus on a replanning scenario for quadrotors where considering time efficiency, non-static initial state and dynamical feasibility is of great significance. We propose a real-time B-spline based kinodynamic (RBK) search algorithm, which transforms a position-only shortest path search (such as A…

Cited by 36SourceScholar