← Search

Shunbo Zhou

34 accepted papers

2026

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

ICML 2026poster

Vision–Language–Action (VLA) models adapt large vision–language backbones to map images and instructions into robot actions. However, prevailing VLAs either generate actions autoregressively in a fixed left-to-right order or attach separate diffusion heads outside the backbone, fragmenting informati…

Cited by 0SourceScholar
2026

Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation

CVPR 2026

Generating realistic hand-object interactions (HOI) videos is a significant challenge due to the difficulty of modeling physical constraints (e.g., contact and occlusion between hands and manipulated objects). Current methods utilize HOI representation as an auxiliary generative objective to guide v

Cited by 0SourceScholar
2026

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes

ICLR 2026poster

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos. However, real-world embodied AI agents—such as robots and self…

Cited by 0SourcecodeScholar
2025

ASCENT: Autonomous Skill Learning Toward Complex Embodied Tasks With Foundation Models

ICRA 2025

Collecting data from simulated scenarios for training robotic skills provides a safer and more controllable alternative to real-world environments. However, it demands considerable effort, including the manual construction of simulation environments, the careful design of tasks, and the challenge of

Cited by 0SourceScholar
2025

ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models

IROS 2025

Recent advancements in Large Language Models (LLMs) have catalyzed numerous efforts to apply these technologies to embodied tasks, with a particular focus on high-level task planning and task decomposition. LLMs face challenges in understanding the physical world, especially regarding spatial, tempo

Cited by 11SourcecodeScholar
2025

Foresee and Act Ahead: Task Prediction and Pre-Scheduling Enabled Efficient Robotic Warehousing

ICRA 2025

In warehousing systems, to enhance efficiency amid surging demand volumes, much attention has been placed on how to reasonably allocate tasks of delivery to robots. However, the labor of robots is still inevitably wasted to some extent. In this paper, we propose a pre-scheduling enhanced warehousing

Cited by 1SourceScholar
2025

GPGS: Geometric Priors for 3D Gaussian Splatting in Structural Environments

IROS 2025

Recently, 3D Gaussian Splatting (3DGS) has garnered significant attention for its remarkable capacity to efficiently synthesize novel views with high fidelity. Nevertheless, 3DGS encounters challenges in accurately representing the geometry of real-world scenes. To address this issue, previous metho

Cited by 0SourceScholar
2025

Graph2Scene: Versatile 3D Indoor Scene Generation with Interaction-aware Scene Graph

IROS 2025

Embodied artificial intelligence requires a wide variety of large-scale simulated environments for development. Previous scene reconstruction approaches based on multiview images can produce high-fidelity 3D scenes but lack diversity. In contrast, existing prompt-based scene generation approaches ca

Cited by 0SourceScholar
2025

Ms. NAMI: Multimodal Semantic Navigation on Relative Metric Intention Graph

ICRA 2025

Embodied navigation in unknown environments presents the significant challenge of integrating tasks with multimodal goals into a unified framework. In this paper, we propose the Multimodal Semantic Navigation on Relative Metric Intention Graph (Ms. NAMI), a framework that integrates various navigati

Cited by 0SourceScholar
2025

OccluGaussian: Occlusion-Aware Gaussian Splatting for Large Scene Reconstruction and Rendering

ICCV 2025poster

In large-scale scene reconstruction using 3D Gaussian splatting, it is common to partition the scene into multiple smaller regions and reconstruct them individually. However, existing division methods are occlusion-agnostic, meaning that each region may contain areas with severe occlusions. As a res…

2025

PanopticSplatting: End-to-End Panoptic Gaussian Splatting

IROS 2025

Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are multi-staged, suffering from the accumulated errors and the d

Cited by 2SourceScholar
2024

Distractor-Free Novel View Synthesis via Exploiting Memorization Effect in Optimization

ECCV 2024poster

"Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have greatly advanced novel view synthesis, which is capable of photo-realistic rendering. However, these methods require the foundational assumption of the static scene (, consistent lighting condition and persistent object positions),…

2024

LoS: Local Structure-Guided Stereo Matching

CVPR 2024poster

Estimating disparities in challenging areas is difficult and limits the performance of stereo matching models. In this paper we exploit local structure information (LSI) to enhance stereo matching. Specifically our LSI comprises a series of key elements including the slant plane (parameterised by di…

Cited by 14SourcePDFScholar
2024

PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction

IROS 2024

Panoptic reconstruction is a challenging task in 3D scene understanding. However, most existing methods heavily rely on pre-trained semantic segmentation models and known 3D object bounding boxes for 3D panoptic segmentation, which is not available for in-the-wild scenes. In this paper, we propose a

Cited by 8SourceScholar
2024

SCALE: Self-Correcting Visual Navigation for Mobile Robots via Anti-Novelty Estimation

ICRA 2024poster

Although visual navigation has been extensively studied using deep reinforcement learning, online learning for real-world robots remains a challenging task. Recent work directly learned from offline dataset to achieve broader generalization in the real-world tasks, which, however, faces the out-of-d…

Cited by 2SourcecodeScholar
2024

Scale Disparity of Instances in Interactive Point Cloud Segmentation

IROS 2024poster

Interactive point cloud segmentation has become a pivotal task for understanding 3D scenes, enabling users to guide segmentation models with simple interactions such as clicks, therefore significantly reducing the effort required to tailor models to diverse scenarios and new categories. However, in…

Cited by 2SourceScholar
2024

Toward Universal and Scalable Road Graph Partitioning for Efficient Multi-Robot Path Planning

IROS 2024

To date, multi-robot path planning has primarily been addressed by centralized solvers, typically aiming to maintain optimality. However, given its NP-hard nature, directly applying existing solvers in large and complex scenarios proves inefficient. A promising alternative lies in adopting a divide-

Cited by 1SourceScholar
2024

Traffic Flow Learning Enhanced Large-Scale Multi-Robot Cooperative Path Planning Under Uncertainties

ICRA 2024poster

Robotic systems with hundreds or even thousands of robots are widely implemented in logistic and industrial applications. In such systems, cooperative path planning is of great importance, as local congestion and motion conflict may greatly degrade system performance, especially in the presence of u…

Cited by 3SourceScholar
2024

iBTC: An Image-Assisting Binary and Triangle Combined Descriptor for Place Recognition by Fusing LiDAR and Camera Measurements

RA-L 2024

In this work, we introduce a novel multimodal descriptor, the image-assisting binary and triangle combined (iBTC) descriptor, which fuses LiDAR (Light Detection and Ranging) and camera measurements for 3D place recognition. The inherent invariance of a triangle to rigid transformations inspires us t

Cited by 8SourceScholar
2023

Audio-Driven High Definetion and Lip-Synchronized Talking Face Generation Based on Face Reenactment

ICASSP 2023accepted

Generating audio-driven photo-realistic talking face has received intensive attention due to its ability to bring more new human-computer interaction experiences. However, previous works struggled to balance high definition, lip synchronization, and low customization costs, which would degrade the u…

Cited by 0SourceScholar
2023

NF-Atlas: Multi-Volume Neural Feature Fields for Large Scale LiDAR Mapping

RA-L 2023

LiDAR Mapping has been a long-standing problem in robotics. Recent progress in neural implicit representation has brought new opportunities to robotic mapping. In this letter, we propose the multi-volume neural feature fields, called NF-Atlas, which bridge the neural feature volumes with pose graph

Cited by 21SourceScholar
2021

Inertial Aided 3D LiDAR SLAM with Hybrid Geometric Primitives in Large-scale Environments

ICRA 2021poster

This paper presents a comprehensive inertial aided 3D LiDAR SLAM system with hybrid geometric primitives in large-scale environments, including a tightly-coupled LiDAR-Inertial-Odometry (LIO), a global mapping module supported by learning-based loop closure detection and a sub-maps matching algorith…

Cited by 7SourceScholar
2020

CUHK-AHU Dataset: Promoting Practical Self-Driving Applications in the Complex Airport Logistics, Hill and Urban Environments

IROS 2020poster

This paper presents a novel dataset targeting three types of challenging environments for autonomous driving, i.e., the industrial logistics environment, the undulating hill environment and the mixed complex urban environment. To the best of the author’s knowledge, similar dataset has not been publi…

Cited by 5SourceScholar
2020

Online Trajectory Planning for an Industrial Tractor Towing Multiple Full Trailers

ICRA 2020poster

This paper presents a novel solution for online trajectory planning of a full-size tractor-trailers vehicle composed of a car-like tractor and arbitrary number of passive full trailers. The motion planning problem for such systems was rarely addressed due to the complex nonlinear dynamics. A simulat…

Cited by 16SourceScholar
2020

Robust Dynamic State Estimation for Lateral Control of an Industrial Tractor Towing Multiple Passive Trailers

IROS 2020poster

In this paper, we propose a dynamic state estimation framework for lateral control of a heavy tractor-trailers system using only mass-produced low-cost sensors. This issue is challenging since the lateral velocity of the lead tractor is difficult to measure directly. The performance of existing dyna…

Cited by 0SourceScholar
2020

Robust Path Following of the Tractor-Trailers System in GPS-Denied Environments

RA-L 2020

This letter reports a general path following framework for the tractor-trailers system in Global Positioning System (GPS)-denied environments. Compared to existing methods, this approach prioritizes a robust, cost-optimized, and easy-to-implement solution. First, to achieve accurate path following,

Cited by 27SourceScholar
2019

A Hierarchical Framework for Coordinating Large-Scale Robot Networks

ICRA 2019poster

In this paper, we study the cooperative path planning and motion coordination problems of the multi-robot system with large number of robots, aiming for practical applications in robotic warehouses and automated transportation systems. Particularly, we solve the life-long planning problem and guaran…

Cited by 15SourceScholar
2019

LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment Analysis

ICCV 2019poster

Point cloud based place recognition is still an open issue due to the difficulty in extracting local features from the raw 3D point cloud and generating the global descriptor, and it's even harder in the large-scale dynamic environments. In this paper, we develop a novel deep neural network, named L…

Cited by 345PDFScholar
2019

Modelling and Dynamic Tracking Control of Industrial Vehicles with Tractor-trailer Structure

IROS 2019poster

Existing works on control of tractor-trailers systems only consider the kinematics model without taking dynamics into account. Also, most of them treat the issue as a pure control theory problem whose solutions are difficult to implement. This paper presents a trajectory tracking control approach fo…

Cited by 21SourceScholar
2019

SeqLPD: Sequence Matching Enhanced Loop-Closure Detection Based on Large-Scale Point Cloud Description for Self-Driving Vehicles

IROS 2019poster

Place recognition and loop-closure detection are main challenges in the localization, mapping and navigation tasks of self-driving vehicles. In this paper, we solve the loop-closure detection problem by incorporating the deep-learning based point cloud description method and the coarse-to-fine seque…

Cited by 70SourceScholar
2019

Vision-Based Dynamic Control of Car-Like Mobile Robots

ICRA 2019poster

Most existing controllers for Car-Like Mobile Robots (CLMR) are designed to handle dynamic effects by decoupling speed and steering controls, also assume that full states are accessible, which are unrealistic for real-world applications. This paper presents a combined speed and steering control syst…

Cited by 9SourceScholar
2018

Vision-Based State Estimation and Trajectory Tracking Control of Car-Like Mobile Robots with Wheel Skidding and Slipping

IROS 2018poster

Most existing trajectory tracking controllers are based on non-skidding and non-slipping assumptions, also assume that full states are accessible, which is unrealistic for real-world applications due to tire-road interaction. This paper presents a novel vision-based approach to achieve high performa…

Cited by 12SourceScholar