← Search

Xinge Zhu

43 accepted papers

2026

From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

ICML 2026poster

Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action. Existing generative policies typically adopt a "Generation-from-Noise" paradig…

Cited by 0SourceScholar
2026

HUMOF: Human Motion Forecasting in Interactive Social Scenes

ICLR 2026poster

Complex dynamic scenes present significant challenges for predicting human behavior due to the abundance of interaction information, such as human-human and human-environment interactions. These factors complicate the analysis and understanding of human behavior, thereby increasing the uncertainty i…

Cited by 0SourceScholar
2026

ReMoGen: Real-time Human Interaction-to-Reaction Generation via Modular Learning from Diverse Data

CVPR 2026

Human behaviors in real-world environments are inherently interactive, with an individual's motion shaped by surrounding agents and the scene. Such capabilities are essential for applications in virtual avatars, interactive animation, and human-robot collaboration. We target real-time human interact

Cited by 0SourceScholar
2025

Can LVLMs Obtain a Driver’s License? A Benchmark Towards Reliable AGI for Autonomous Driving

AAAI 2025technical

Large Vision-Language Models (LVLMs) have recently garnered significant attention, with many efforts aimed at harnessing their general knowledge to enhance the interpretability and robustness of autonomous driving models. However, LVLMs typically rely on large, general-purpose datasets and lack the…

Cited by 4SourcePDFScholar
2025

EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment

ICCV 2025poster

Dexterous robotic hands often struggle to generalize effectively in complex environments due to models trained on low-diversity data. However, the real world presents an inherently unbounded range of scenarios. A natural solution is to enable robots learning from experience in complex environments--…

Cited by 0SourcePDFScholar
2025

FreeCap: Hybrid Calibration-Free Motion Capture in Open Environments

AAAI 2025technical

We propose a novel hybrid calibration-free method FreeCap to accurately capture global multi-person motions in open environments. Our system combines a single LiDAR with expandable moving cameras, allowing for flexible and precise motion estimation in a unified world coordinate. In particular, We in…

Cited by 0SourcePDFScholar
2025

FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens

NeurIPS 2025poster

Learning effective visuomotor policies for robotic manipulation is challenging, as it requires generating precise actions while maintaining computational efficiency. Existing methods remain unsatisfactory due to inherent limitations in the essential action representation and the basic network archit…

Cited by 0SourcecodeScholar
2025

ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving

ICCV 2025poster

End-to-end autonomous driving has emerged as a promising approach to unify perception, prediction, and planning within a single framework, reducing information loss and improving adaptability. However, existing methods often rely on fixed and sparse trajectory supervision, limiting their ability to…

Cited by 0SourcePDFScholar
2025

STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation

IROS 2025

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature misalignment due to inadequate decoupling of spatio-temporal dynami

Cited by 5SourcecodeScholar
2024

A Unified Framework for Human-centric Point Cloud Video Understanding

CVPR 2024poster

Human-centric Point Cloud Video Understanding (PVU) is an emerging field focused on extracting and interpreting human-related features from sequences of human point clouds further advancing downstream human-centric tasks and applications. Previous works usually focus on tackling one specific task an…

Cited by 2SourcePDFScholar
2024

HUNTER: Unsupervised Human-centric 3D Detection via Transferring Knowledge from Synthetic Instances to Real Scenes

CVPR 2024poster

Human-centric 3D scene understanding has recently drawn increasing attention driven by its critical impact on robotics. However human-centric real-life scenarios are extremely diverse and complicated and humans have intricate motions and interactions. With limited labeled data supervised methods are…

Cited by 3SourcePDFScholar
2024

OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree Queries

NeurIPS 2024poster

Occupancy prediction has increasingly garnered attention in recent years for its fine-grained understanding of 3D scenes. Traditional approaches typically rely on dense, regular grid representations, which often leads to excessive computational demands and a loss of spatial details for small objects…

2024

TASeg: Temporal Aggregation Network for LiDAR Semantic Segmentation

CVPR 2024poster

Training deep models for LiDAR semantic segmentation is challenging due to the inherent sparsity of point clouds. Utilizing temporal data is a natural remedy against the sparsity problem as it makes the input signal denser. However previous multi-frame fusion algorithms fall short in utilizing suffi…

2024

WildRefer: 3D Object Localization in Large-scale Dynamic Scenes with Multi-modal Visual Data and Natural Language

ECCV 2024poster

"We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method, dubbed WildRefer, for this task by fully utilizing the rich appe…

2023

CLIP2Scene: Towards Label-Efficient 3D Scene Understanding by CLIP

CVPR 2023poster

Contrastive Language-Image Pre-training (CLIP) achieves promising results in 2D zero-shot and few-shot learning. Despite the impressive performance in 2D, applying CLIP to help the learning in 3D scene understanding has yet to be explored. In this paper, we make the first attempt to investigate how…

2023

ContrastMotion: Self-supervised Scene Motion Learning for Large-Scale LiDAR Point Clouds

IJCAI 2023poster

In this paper, we propose a novel self-supervised motion estimator for LiDAR-based autonomous driving via BEV representation. Different from usually adopted self-supervised strategies for data-level structure consistency, we predict scene motion via feature-level consistency between pillars in conse…

2023

GANet: Goal Area Network for Motion Forecasting

ICRA 2023poster

Predicting the future motion of road participants is crucial for autonomous driving but is extremely challenging due to staggering motion uncertainty. Recently, most motion forecasting methods resort to the goal-based strategy, i.e., predicting endpoints of motion trajectories as conditions to regre…

Cited by 89SourcecodeScholar
2023

Human-centric Scene Understanding for 3D Large-scale Scenarios

ICCV 2023poster

Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset…

Cited by 26PDFcodeScholar
2023

One Training for Multiple Deployments: Polar-based Adaptive BEV Perception for Autonomous Driving

ICRA 2023poster

Current on-board chips usually have different computing power, which means multiple training processes are needed for adapting the same learning-based algorithm to different chips, costing huge computing resources. The situation becomes even worse for 3D perception methods with large models. Previou…

Cited by 5SourceScholar
2023

PARTNER: Level up the Polar Representation for LiDAR 3D Object Detection

ICCV 2023poster

Recently, polar-based representation has shown promising properties in perceptual tasks. In addition to Cartesian-based approaches, which separate point clouds unevenly, representing point clouds as polar grids has been recognized as an alternative due to (1) its advantage in robust performance unde…

Cited by 10PDFcodeScholar
2023

Rethinking Range View Representation for LiDAR Segmentation

ICCV 2023poster

LiDAR segmentation is crucial for autonomous driving perception. Recent trends favor point- or voxel-based methods as they often yield better performance than the traditional range view representation. In this work, we unveil several key factors in building powerful range view models. We observe tha…

Cited by 173PDFScholar
2023

SCPNet: Semantic Scene Completion on Point Cloud

CVPR 2023highlight

Training deep models for semantic scene completion is challenging due to the sparse and incomplete input, a large quantity of objects of diverse scales as well as the inherent label noise for moving objects. To address the above-mentioned problems, we propose the following three solutions: 1) Redesi…

Cited by 95SourcePDFScholar
2023

See More and Know More: Zero-shot Point Cloud Segmentation via Multi-modal Visual Data

ICCV 2023poster

Zero-shot point cloud segmentation aims to make deep models capable of recognizing novel objects in point cloud that are unseen in the training phase. Recent trends favor the pipeline which transfers knowledge from seen classes with labels to unseen classes without labels. They typically align visua…

Cited by 33PDFScholar
2023

Towards Label-free Scene Understanding by Vision Foundation Models

NeurIPS 2023poster

Vision foundation models such as Contrastive Vision-Language Pre-training (CLIP) and Segment Anything (SAM) have demonstrated impressive zero-shot performance on image classification and segmentation tasks. However, the incorporation of CLIP and SAM for label-free scene understanding has yet to be e…

2023

UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg Codebase

ICCV 2023poster

Point-, voxel-, and range-views are three representative forms of point clouds. All of them have accurate 3D measurements but lack color and texture information. RGB images are a natural complement to these point cloud views and fully utilizing the comprehensive information of them benefits more rob…

Cited by 46PDFcodeScholar
2022

Point-to-Voxel Knowledge Distillation for LiDAR Semantic Segmentation

CVPR 2022poster

This article addresses the problem of distilling knowledge from a large teacher model to a slim student network for LiDAR semantic segmentation. Directly employing previous distillation approaches yields inferior results due to the intrinsic challenges of point cloud, i.e., sparsity, randomness and…

Cited by 215PDFcodeScholar
2022

STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded Scenes

CVPR 2022poster

Accurately detecting and tracking pedestrians in 3D space is challenging due to large variations in rotations, poses and scales. The situation becomes even worse for dense crowds with severe occlusions. However, existing benchmarks either only provide 2D annotations, or have limited 3D annotations w…

Cited by 49PDFcodeScholar
2022

TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection With Transformers

CVPR 2022poster

LiDAR and camera are two important sensors for 3D object detection in autonomous driving. Despite the increasing popularity of sensor fusion in this field, the robustness against inferior image conditions, e.g., bad illumination and sensor misalignment, is under-explored. Existing fusion methods are…

Cited by 806PDFcodeScholar
2021

AdaStereo: A Simple and Efficient Approach for Adaptive Stereo Matching

CVPR 2021poster

Recently, records on stereo matching benchmarks are constantly broken by end-to-end disparity networks. However, the domain adaptation ability of these deep models is quite poor. Addressing such problem, we present a novel domain-adaptive pipeline called AdaStereo that aims to align multi-level repr…

Cited by 91PDFScholar
2021

Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR Segmentation

CVPR 2021poster

State-of-the-art methods for large-scale driving-scene LiDAR segmentation often project the point clouds to 2D space and then process them via 2D convolution. Although this corporation shows the competitiveness in the point cloud, it inevitably alters and abandons the 3D topology and geometric relat…

Cited by 675PDFcodeScholar
2021

LiDAR-Based Panoptic Segmentation via Dynamic Shifting Network

CVPR 2021poster

With the rapid advances of autonomous driving, it becomes critical to equip its sensing system with more holistic 3D perception. However, existing works focus on parsing either the objects (e.g. cars and pedestrians) or scenes (e.g. trees and buildings) from the LiDAR sensor. In this work, we addres…

Cited by 114PDFcodeScholar
2021

Probabilistic and Geometric Depth: Detecting Objects in Perspective

CoRL 2021poster

3D object detection is an important capability needed in various practical applications such as driver assistance systems. Monocular 3D detection, a representative general setting among image-based approaches, provides a more economical solution than conventional settings relying on LiDARs but still…

Cited by 325SourcecodeScholar
2020

AutoTrajectory: Label-free Trajectory Extraction and Prediction from Videos using Dynamic Points

ECCV 2020poster

Current methods for trajectory prediction operate in supervised manners, and therefore require vast quantities of corresponding ground truth data for training. In this paper, we present a novel, label-free algorithm, AutoTrajectory, for trajectory extraction and prediction to use raw videos directly…

2020

SelfVoxeLO: Self-supervised LiDAR Odometry with Voxel-based Deep Neural Networks

CoRL 2020

Recent learning-based LiDAR odometry methods have demonstrated their competitiveness. However, most methods still face two substantial challenges: 1) the 2D projection representation of LiDAR data cannot effectively encode 3D structures from the point clouds; 2) the needs for a large amount of label

2020

Tensor Low-Rank Reconstruction for Semantic Segmentation

ECCV 2020poster

Context information plays an indispensable role in the success of semantic segmentation. Recently, non-local self-attention based methods are proved to be effective for context information collection. Since desired context consists of spatial-wise and channel-wise attentions, the 3D representation i…

Cited by 88SourcePDFScholar
2019

Adapting Object Detectors via Selective Cross-Domain Alignment

CVPR 2019poster

State-of-the-art object detectors are usually trained on public datasets. They often face substantial difficulties when applied to a different domain, where the imaging condition differs significantly and the corresponding annotated data are unavailable (or expensive to acquire). A natural remedy is…

Cited by 436PDFcodeScholar
2019

Depth Completion From Sparse LiDAR Data With Depth-Normal Constraints

ICCV 2019poster

Depth completion aims to recover dense depth maps from sparse depth measurements. It is of increasing importance for autonomous driving and draws increasing attention from the vision community. Most of the current competitive methods directly train a network to learn a mapping from sparse depth inpu…

Cited by 254PDFScholar
2019

Not All Areas Are Equal: Transfer Learning for Semantic Segmentation via Hierarchical Region Selection

CVPR 2019oral

The success of deep neural networks for semantic segmentation heavily relies on large-scale and well-labeled datasets, which are hard to collect in practice. Synthetic data offers an alternative to obtain ground-truth labels for free. However, models directly trained on synthetic data often struggle…

Cited by 89PDFScholar
2018

Penalizing Top Performers: Conservative Loss for Semantic Segmentation Adaptation

ECCV 2018poster

Due to the expensive and time-consuming annotations (e.g., segmentation) for real-world images, recent works in computer vision resort to synthetic data. However, the performance on the real image often drops significantly because of the domain shift between the synthetic data and the real images. I…

Cited by 135SourcePDFScholar