← Search

Zhongxue Gan

41 accepted papers

2026

CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning

CVPR 2026

Multimodal learning aims to capture both shared and private information from multiple modalities. However, existing methods that project all modalities into a single latent space for fusion often overlook the asynchronous, multi-level semantic structure of multimodal data. This oversight induces sem

Cited by 0SourceScholar
2026

CMoE: Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots

ICRA 2026poster

For effective deployment in real-world environments, humanoid robots must autonomously navigate a diverse range of complex terrains with abrupt transitions. While the Vanilla mixture of experts (MoE) framework is theoretically capable of modeling diverse terrain features, in practice, the gating net…

2026

Collaborative Learning of Local 3D Occupancy Prediction and Versatile Global Occupancy Mapping

ICRA 2026poster

Vision-based 3D semantic occupancy prediction is vital for autonomous driving, enabling unified modeling of static infrastructure and dynamic agents. Global occupancy maps serve as long-term memory priors, providing valuable historical context that enhances local perception. This is particularly imp…

2026

Drive in Corridors: Enhancing the Safety of End-To-End Autonomous Driving Via Corridor Learning and Planning

ICRA 2026poster

Safety remains one of the most critical challenges in autonomous driving systems. In recent years, the end-to-end driving has shown great promise in advancing vehicle autonomy in a scalable manner. However, existing approaches often face safety risks due to the lack of explicit behavior constraints.…

2026

GIL-3D: U-Shaped Diffusion Transformers for Generalizable 3D Imitation Learning

RA-L 2026

Imitation learning with 3D vision effectively alleviates the impact of variations in lighting, background, and texture. It exhibits superior robustness compared to 2D-based methods. However, existing 3D imitation learning methods often suffer from performance degradation as the task horizon increase

Cited by 0SourceScholar
2026

Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration

ICML 2026poster

Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers from two limitations: modality dominance, where optimization gravitates towards the path of least resistance, ignoring …

Cited by 0SourceScholar
2026

Learning a Unified Risk Map for Autonomous Driving in Partially Observable Environments

RA-L 2026

Occlusion-aware prediction remains a critical challenge in autonomous driving due to the inherent uncertainty of unobserved regions. Existing approaches either overestimate risk based on reachable states or struggle to predict accurate trajectories under high occlusion uncertainty. To address these

Cited by 0SourceScholar
2026

OccLLaMA: A Unified Occupancy-Language-Action World Model for Enhancing Motion Planning Via Multi-Task Learning

ICRA 2026poster

Scene understanding via multi-modal large language models and scene forecasting with world models have advanced the development of autonomous driving. The former maps visual inputs to driving-specific outputs, neglecting spatial reasoning and world dynamics. The latter captures world dynamics, lacki…

Cited by 0codeScholar
2026

The Folding Hand: Anthropomorphic Robotic Hands with a Compact Reconfigurable Humanoid Palm Design

ICRA 2026poster

The human palm is a remarkable and highly functional part of the hand that significantly contributes to dexterity, grasp versatility, and overall manipulation capability. The metacarpophalangeal joints (MCP) of the palm facilitate movement of the fingers for flexion, extension, abduction, adduction,…

Cited by 0SourceScholar
2026

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

CVPR 2026

Multimodal Sentiment Analysis (MSA) integrates language, visual, and acoustic modalities to infer human sentiment. Most existing methods either focus on globally shared representations or modality-specific features, while overlooking signals that are shared only by certain modality pairs. This limit

Cited by 0SourceScholar
2026

Unveiling the Surprising Efficacy of Navigation Understanding in End-To-End Autonomous Driving

ICRA 2026poster

Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-end autonomous driving systems tend to over-rely on local scene understanding while failing to utilize global navigation i…

2026

VINGS-Mono: Visual-Inertial Gaussian Splatting Monocular SLAM in Large Scenes

ICRA 2026poster

VINGS-Mono is a monocular inertial Gaussian Splatting (GS) SLAM framework designed for large-scale scenes. It integrates four main components: VIO Front End, 2D Gaussian Map, NVS Loop Closure, and Dynamic Eraser. The VIO Front End processes RGB frames with dense bundle adjustment and uncertainty est…

2025

A Modified Resistance Model for Magnetic Honeycomb Robots to Navigate in Low Reynolds Number Fluids

ICRA 2025

In recent years, magnetically controlled microrobots have garnered significant attention. This paper presents the H-robot, a self-designed microrobot featuring an innovative structure. The H-robot features a honeycomb porous spherical design specifically engineered to enhance cargo capacity. A new d

Cited by 0SourceScholar
2025

Drive in Corridors: Enhancing the Safety of End-to-End Autonomous Driving via Corridor Learning and Planning

RA-L 2025

Safety remains one of the most critical challenges in autonomous driving systems. In recent years, the end-to-end driving has shown great promise in advancing vehicle autonomy in a scalable manner. However, existing approaches often face safety risks due to the lack of explicit behavior constraints.

Cited by 3SourcecodeScholar
2025

HGS-Planner: Hierarchical Planning Framework for Active Scene Reconstruction Using 3D Gaussian Splatting

ICRA 2025

In complex missions such as search and rescue, robots must make intelligent decisions in unknown environments, relying on their ability to perceive and understand their surroundings. High-quality and real-time reconstruction enhances situational awareness and is crucial for intelligent robotics. Tra

Cited by 19SourceScholar
2025

Once-Tuning-Multiple-Variants: Tuning Once and Expanded as Multiple Vision-Language Model Variants

CVPR 2025poster

Vision-language model (VLM) is one of the most important models for multi-modal tasks. Real industrial applications often meet the challenge of adapting VLMs to different scenarios, such as varying hardware platforms or performance requirements. Traditional methods involve training or fine-tuning to…

2025

Spherical Scissor-Like Reconfigurable Palm Design in Robotic Hands: Insights from Human Hand Functionality

IROS 2025

The human palm demonstrates spatial reconfigurability during the gripping process and forms a spherical grasping envelope. Based on these observations, this study designs a reconfigurable spherical palm that incorporates a spatial scissor mechanism, which only requires a single actuator to reshape t

Cited by 0SourceScholar
2025

The Folding Hand: Anthropomorphic Robotic Hands With a Compact Reconfigurable Humanoid Palm Design

RA-L 2025

The human palm is a remarkable and highly functional part of the hand that significantly contributes to dexterity, grasp versatility, and overall manipulation capability. The metacarpophalangeal joints (MCP) of the palm facilitate movement of the fingers for flexion, extension, abduction, adduction,

Cited by 1SourceScholar
2025

Topology-Driven Trajectory Optimization for Modelling Controllable Interactions Within Multi-Vehicle Scenario

IROS 2025

Trajectory optimization in multi-vehicle scenarios faces challenges due to its non-linear, non-convex properties and sensitivity to initial values, making interactions between vehicles difficult to control. In this paper, inspired by topological planning, we propose a differentiable local homotopy i

Cited by 0SourceScholar
2025

UAV-DETR: Efficient End-to-End Object Detection for Unmanned Aerial Vehicle Imagery

IROS 2025

Unmanned aerial vehicle object detection (UAV-OD) has been widely used in various scenarios. However, most existing UAV-OD algorithms rely on manually designed components, which require extensive tuning. End-to-end models that do not depend on such manually designed components are mainly designed fo

Cited by 54SourcecodeScholar
2024

A Soft Continuum Robot With Self-Controllable Variable Curvature

RA-L 2024

This letter introduces a new type of soft continuum robot, called SCoReS, which is capable of self-controlling continuously its curvature at the segment level; in contrast to previous designs which either require external forces or machine elements, or whose variable curvature capabilities are discr

Cited by 12SourceScholar
2024

CenterCoop: Center-Based Feature Aggregation for Communication-Efficient Vehicle-Infrastructure Cooperative 3D Object Detection

RA-L 2024

Vehicle-Infrastructure Cooperative (VIC) 3D object detection is a challenging task for balancing communication bandwidth and detection performance. Intermediate fusion is recently studied to reach a better balance by transferring feature maps. Existing works mainly perform spatial-wise fusion and ad

Cited by 10SourceScholar
2024

HGS-Mapping: Online Dense Mapping Using Hybrid Gaussian Representation in Urban Scenes

RA-L 2024

Online dense mapping of urban scenes forms a fundamental cornerstone for scene understanding and navigation of autonomous vehicles. Recent advancements in dense mapping methods are mainly based on NeRF, whose rendering speed is too slow to meet online requirements. 3D Gaussian Splatting (3DGS), with

Cited by 20SourceScholar
2024

OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data

ICRA 2024poster

In the era of big data and large models, automatic annotating functions for multi-modal data are of great significance for real-world AI-driven applications, such as autonomous driving and embodied AI. Unlike traditional closed-set annotation, open-vocabulary annotation is essential to achieve human…

Cited by 13SourcecodeScholar
2024

Salpot: A Jet Propulsion Swimmer With Scissor Structure and Bilateral Apertures

RA-L 2024

In recent years, researchers have increasingly turned to marine organisms for inspiration in designing underwater robots. While most robots rely on jet propulsion, akin to squid or jellyfish, using a single posterior aperture for water intake and expulsion, there are few incorporating an additional

Cited by 6SourceScholar
2024

Spear: Evaluate the Adversarial Robustness of Compressed Neural Models

IJCAI 2024poster

As Artificial Intelligence evolves, the neural models vulnerable to adversarial attacks may produce fatal results in critical applications. This paper mainly discusses the robustness of the compressed neural models facing adversarial attacks. A few studies discuss the interaction between model compr…

2024

Swift-Mapping: Online Neural Implicit Dense Mapping in Urban Scenes

AAAI 2024technical

Online dense mapping of urban scenes is of paramount importance for scene understanding of autonomous navigation. Traditional online dense mapping methods fuse sensor measurements (vision, lidar, etc.) across time and space via explicit geometric correspondence. Recently, NeRF-based methods have pro…

Cited by 2SourcePDFScholar
2023

Adversarial Amendment is the Only Force Capable of Transforming an Enemy into a Friend

IJCAI 2023poster

Adversarial attack is commonly regarded as a huge threat to neural networks because of misleading behavior. This paper presents an opposite perspective: adversarial attacks can be harnessed to improve neural models if amended correctly. Unlike traditional adversarial defense or adversarial training…

2023

Boost Transformer-based Language Models with GPU-Friendly Sparsity and Quantization

ACL 2023findings

Along with the performance improvement in NLP domain, the sizes of transformer-based language models (TLM) are also dramatically increased. Some prior works intend to compress TLM models into more compact forms, but do not fully consider the hardware characters may not support the efficient executio…

2023

Boost Vision Transformer With GPU-Friendly Sparsity and Quantization

CVPR 2023poster

The transformer extends its success from the language to the vision domain. Because of the numerous stacked self-attention and cross-attention blocks in the transformer, which involve many high-dimensional tensor multiplication operations, the acceleration deployment of vision transformer on GPU har…

2023

FlowMap: Path Generation for Automated Vehicles in Open Space Using Traffic Flow

ICRA 2023poster

There is extensive literature on perceiving road structures by fusing various sensor inputs such as lidar point clouds and camera images using deep neural nets. Leveraging the latest advance of neural architects (such as transformers) and bird-eye-view (BEV) representation, the road cognition accura…

Cited by 4SourceScholar
2023

Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement Augmentation

ICCV 2023poster

Learning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficie…

Cited by 25PDFScholar
2023

Mechanical Intelligence for Prehensile In-Hand Manipulation of Spatial Trajectories

ICRA 2023poster

The application of mechanical and other physical properties to the development of robotic systems that can easily adapt to changing external situations is known as mechanical intelligence. Following this concept, many robot hand designs can produce self-adaptive and versatile grasps with simple unde…

Cited by 2SourceScholar
2022

Efficient Universal Shuffle Attack for Visual Object Tracking

ICASSP 2022accepted

Recently, adversarial attacks have been applied in visual object tracking to deceive deep trackers by injecting imperceptible perturbations into video frames. However, previous work only generates the video-specific perturbations, which restricts its application scenarios. In addition, existing atta…

Cited by 0SourceScholar
2022

Learning From Demonstrations Via Multi-Level and Multi-Attention Domain-Adaptive Meta-Learning

RA-L 2022

Despite significant advances in few-shot classification, object detection, or speech recognition in recent years, training an effective robot to adapt to previously unseen environments in a small data regime is still a long-lasting problem for learning from demonstrations (LfD). A promising solution

Cited by 3SourceScholar
2022

Learning With Dual Demonstration Domains: Random Domain-Adaptive Meta-Learning

RA-L 2022

Although robots have been widely applied in various fields, allowing a robot to perform a wide range of tasks like humans is a significant challenge. One promising method is meta-learning, which enables robots to learn from demonstrations with the concept of “learning to learn.” Howeve

Cited by 7SourceScholar
2020

Multi-Scale Deep Feature Fusion for Vehicle Re-Identification

ICASSP 2020accepted

Vehicle re-identification (re-id) is challenging due to the small inter-class distance. The differences between similar vehicles can be extremely subtle and only captured at particular scales and semantic levels. In this paper, we propose a novel Multi-Scale Deep Feature Fusion Network (MSDeep) to c…

Cited by 0SourceScholar