← Search

Zheng Zhu

63 accepted papers

2026

Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding

ICLR 2026poster

Text-to-Motion generation has become a fundamental task in human-machine interaction, enabling the synthesis of realistic human motions from natural language descriptions. Although recent advances in large language models and reinforcement learning have contributed to high-quality motion generation,…

Cited by 0SourceScholar
2026

ORV: 4D Occupancy-centric Robot Video Generation

CVPR 2026

Recent embodied intelligence suffers from data scarcity, while conventional simulators lack visual realism. Controllable video generation is emerging as a promising data engine, yet current action-conditioned methods still fall short: generated videos are limited in fidelity and temporal consistency

Cited by 0SourcecodeScholar
2026

R2RGen: Real-to-Real 3D Data Generation for Spatially-generalized Robotic Manipulation

RSS 2026poster

Towards the aim of generalized robotic manipulation, spatial generalization is the most fundamental capability that requires the policy to work robustly under different spatial distribution of objects, environment and agent itself. To achieve this, substantial human demonstrations need to be collect…

Cited by 0SourceScholar
2026

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization

ICLR 2026poster

Vector quantization (VQ) is a key component in discrete tokenizers for image generation, but its training is often unstable due to straight-through estimation bias, one-step-behind updates, and sparse codebook gradients, which lead to suboptimal reconstruction performance and low codebook usage. In…

Cited by 0SourceScholar
2026

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

ICML 2026poster

Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a severe performance collapse. We attribute this degradation to the spatially agnostic nature of existing reduction methods, …

Cited by 0SourceScholar
2026

SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead

CVPR 2026

Vision-Language-Action (VLA) models built on pretrained Vision-Language Models (VLMs) show strong potential but are limited in practicality due to their large parameter counts. To mitigate this issue, using a lightweight VLM has been explored, but it compromises spatiotemporal reasoning. Although so

Cited by 0SourcecodeScholar
2025

Adjacent-view Transformers for Supervised Surround-view Depth Estimation

IROS 2025

Depth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monocular depth estimation in the past decades, these attempts are mainly conducted on the KITTI benchmark with only front-vi

Cited by 5SourcecodeScholar
2025

DetRF: Detachable Novel Views Synthesis of Dynamic Scenes Using Backdrop-Driven Neural Radiance Fields

AAAI 2025technical

Representing and synthesizing novel views in real-world dynamic scenes from casual monocular videos is a long-standing problem. Existing solutions typically approach dynamic scenes by applying geometry techniques or utilizing temporal information between several adjacent frames without considering t…

Cited by 0SourcePDFScholar
2025

DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

AAAI 2025technical

World models have demonstrated superiority in autonomous driving, particularly in the generation of multi-view driving videos. However, significant challenges still exist in generating customized driving videos. In this paper, we propose DriveDreamer-2, which incorporates a Large Language Model (LLM…

Cited by 62SourcePDFScholar
2025

DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation

CVPR 2025poster

Closed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data distributions, which are largely confined to forward-driving scenarios. Conseque…

Cited by 23SourcePDFScholar
2025

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Videos Generation

NeurIPS 2025poster

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant potential for enhancing applications in virtual reality, aug…

Cited by 0SourceScholar
2025

HumanDreamer: Generating Controllable Human-Motion Videos via Decoupled Generation

CVPR 2025poster

Human-motion video generation has been a challenging task, primarily due to the difficulty inherent in learning human body movements. While some approaches have attempted to drive human-centric video generation explicitly through pose control, these methods typically rely on poses derived from exist…

Cited by 2SourcePDFScholar
2025

JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems

CVPR 2025poster

Unmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we int…

Cited by 0SourcePDFScholar
2025

Pruning-Robust Mamba with Asymmetric Multi-Scale Scanning Paths

NeurIPS 2025poster

Mamba has proven efficient for long-sequence modeling in vision tasks. However, when token reduction techniques are applied to improve efficiency, Mamba-based models exhibit drastic performance degradation compared to Vision Transformers (ViTs). This decline is potentially attributed to Mamba's cha…

Cited by 0SourceScholar
2025

ReconDreamer++: Harmonizing Generative and Reconstructive Models for Driving Scene Representation

ICCV 2025poster

Combining reconstruction models with generative models has emerged as a promising paradigm for closed-loop simulation in autonomous driving. For example, ReconDreamer has demonstrated remarkable success in rendering large-scale maneuvers. However, a significant gap remains between the generated data…

Cited by 0SourcePDFScholar
2025

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration

CVPR 2025poster

Closed-loop simulation is crucial for end-to-end autonomous driving. Existing sensor simulation methods (e.g., NeRF and 3DGS) reconstruct driving scenes based on conditions that closely mirror training data distributions. However, these methods struggle with rendering novel trajectories, such as lan…

Cited by 11SourcePDFScholar
2025

WonderTurbo: Generating Interactive 3D World in 0.72 Seconds

ICCV 2025poster

Interactive 3D generation is gaining momentum and capturing extensive attention for its potential to create immersive virtual experiences. However, a critical challenge in current 3D generation technologies lies in achieving real-time interactivity. To address this issue, we introduce WonderTurbo, t…

Cited by 0SourcePDFScholar
2024

DiffBEV: Conditional Diffusion Model for Bird’s Eye View Perception

AAAI 2024technical

BEV perception is of great importance in the field of autonomous driving, serving as the cornerstone of planning, controlling, and motion prediction. The quality of the BEV feature highly affects the performance of BEV perception. However, taking the noises in camera parameters and LiDAR scans into…

2024

DiffusionDepth: Diffusion Denoising Approach for Monocular Depth Estimation

ECCV 2024poster

"Monocular depth estimation is a challenging task that predicts the pixel-wise depth from a single 2D image. Current methods typically model this problem as a regression or classification task. We propose DiffusionDepth, a new approach that reformulates monocular depth estimation as a denoising diff…

2024

DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

ECCV 2024poster

"World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the generation of high-quality driving videos, and driving policies for safe maneuvering…

Cited by 183SourcePDFScholar
2024

Foot Vision: A Vision-Based Multi-Functional Sensorized Foot for Quadruped Robots

RA-L 2024

Quadruped robots equipped with sensorless feet can only provide limited information regarding the foot interaction with the surroundings, limiting their off-road exploration capability in unstructured environments. To tackle this problem, we present Foot Vision, an innovative vision-based sensorized

Cited by 10SourceScholar
2024

One at a Time: Progressive Multi-Step Volumetric Probability Learning for Reliable 3D Scene Perception

AAAI 2024technical

Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes directly with geometric correspondence, attempting to fully address…

Cited by 3SourcePDFScholar
2024

OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models

ECCV 2024poster

"Panoptic Scene Graph Generation (PSG) aims to segment objects and recognize their relations, enabling the structured understanding of an image. Previous methods focus on predicting predefined object and relation categories, hence limiting their applications in the open world scenarios. With the rap…

2024

TAIL: A Terrain-Aware Multi-Modal SLAM Dataset for Robot Locomotion in Deformable Granular Environments

RA-L 2024

Terrain-aware perception holds the potential to improve the robustness and accuracy of autonomous robot navigation in the wilds, thereby facilitating effective off-road traversals. However, the lack of multi-modal perception across various motion patterns hinders the solutions of Simultaneous Locali

Cited by 14SourcecodeScholar
2024

Unified Single-Stage Transformer Network for Efficient RGB-T Tracking

IJCAI 2024poster

Most existing RGB-T tracking networks extract modality features in a separate manner, which lacks interaction and mutual guidance between modalities. This limits the network's ability to adapt to the diverse dual-modality appearances of targets and the dynamic relationships between the modalities. A…

2023

A Simple Baseline for Multi-Camera 3D Object Detection

AAAI 2023technical

3D object detection with surrounding cameras has been a promising direction for autonomous driving. In this paper, we present SimMOD, a Simple baseline for Multi-camera Object Detection, to solve the problem. To incorporate multiview information as well as build upon previous efforts on monocular 3D…

2023

Are We Ready for Vision-Centric Driving Streaming Perception? The ASAP Benchmark

CVPR 2023poster

In recent years, vision-centric perception has flourished in various autonomous driving tasks, including 3D detection, semantic map construction, motion forecasting, and depth estimation. Nevertheless, the latency of vision-centric approaches is too high for practical deployment (e.g., most camera-b…

2023

CompletionFormer: Depth Completion With Convolutions and Vision Transformers

CVPR 2023poster

Given sparse depths and the corresponding RGB images, depth completion aims at spatially propagating the sparse measurements throughout the whole image to get a dense depth prediction. Despite the tremendous progress of deep-learning-based depth completion methods, the locality of the convolutional…

2023

Crafting Monocular Cues and Velocity Guidance for Self-Supervised Multi-Frame Depth Learning

AAAI 2023technical

Self-supervised monocular methods can efficiently learn depth information of weakly textured surfaces or reflective objects. However, the depth accuracy is limited due to the inherent ambiguity in monocular geometric modeling. In contrast, multi-frame depth estimation methods improve depth accuracy…

2023

DREAM: Efficient Dataset Distillation by Representative Matching

ICCV 2023poster

Dataset distillation aims to synthesize small datasets with little information loss from original large-scale ones for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample synthesis process by matching synthetic images and the original ones regarding gradi…

Cited by 125PDFcodeScholar
2023

Divide to Adapt: Mitigating Confirmation Bias for Domain Adaptation of Black-Box Predictors

ICLR 2023top-25%

Domain Adaptation of Black-box Predictors (DABP) aims to learn a model on an unlabeled target domain supervised by a black-box predictor trained on a source domain. It does not require access to both the source-domain data and the predictor parameters, thus addressing the data privacy and portabilit…

2023

DyGait: Exploiting Dynamic Representations for High-performance Gait Recognition

ICCV 2023poster

Gait recognition is a biometric technology that recognizes the identity of humans through their walking patterns. Compared with other biometric technologies, gait recognition is more difficult to disguise and can be applied to the condition of long-distance without the cooperation of subjects. Thus,…

Cited by 48PDFScholar
2023

Efficient and Hybrid Decoder for Local Map Construction in Bird'-Eye-View

ICRA 2023poster

High-definition maps are crucial perception elements for autonomous robot navigation systems, which can provide accurate scene layout and environment information for downstream motion prediction and planning control tasks. Traditional methods based on manual annotation or SLAM algorithms require mas…

Cited by 1SourceScholar
2023

HFT: Lifting Perspective Representations via Hybrid Feature Transformation for BEV Perception

ICRA 2023poster

Restoring an accurate Bird's Eye View (BEV) map plays a crucial role in the perception of autonomous driving. The existing works of lifting representations from frontal view to BEV can be classified into two categories, i.e., Camera model-Based Feature Transformation (CBFT) and Camera model-Free Fea…

Cited by 11SourceScholar
2023

OPERA: Omni-Supervised Representation Learning with Hierarchical Supervisions

ICCV 2023poster

The pretrain-finetune paradigm in modern computer vision facilitates the success of self-supervised learning, which tends to achieve better transferability than supervised learning. However, with the availability of massive labeled data, a natural question emerges: how to train a better model with b…

Cited by 8PDFcodeScholar
2023

OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy Prediction

ICCV 2023poster

The vision-based perception for autonomous driving has undergone a transformation from the bird-eye-view (BEV) representations to the 3D semantic occupancy. Compared with the BEV planes, the 3D semantic occupancy further provides structural information along the vertical direction. This paper presen…

Cited by 207PDFcodeScholar
2023

OpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy Perception

ICCV 2023poster

Semantic occupancy perception is essential for autonomous driving, as automated vehicles require a fine-grained perception of the 3D urban structures. However, existing relevant benchmarks lack diversity in urban scenes, and they only evaluate front-view predictions. Towards a comprehensive benchmar…

Cited by 177PDFcodeScholar
2023

SurroundOcc: Multi-camera 3D Occupancy Prediction for Autonomous Driving

ICCV 2023poster

3D scene understanding plays a vital role in vision-based autonomous driving. While most existing methods focus on 3D object detection, they have difficulty describing real-world objects of arbitrary shapes and infinite classes. Towards a more comprehensive perception of a 3D scene, in this paper, w…

Cited by 260PDFcodeScholar
2023

Wheel Vision: Wheel-Terrain Interaction Measurement and Analysis Using a Sensorized Transparent Wheel on Deformable Terrains

RA-L 2023

The off-road locomotion of wheeled mobile robots (WMRs) over soft terrains can be quite challenging due to the complicated wheel-terrain interaction (WTI). To avoid unforeseen non-geometric hazards such as excessive sinkage or slippage, it is crucial to oversee these terrain-related uncertainties. H

Cited by 13SourceScholar
2023

Wheel-Terrain Contact Geometry Estimation and Interaction Analysis Using Aside-Wheel Camera Over Deformable Terrains

RA-L 2023

Wheeled mobile robots (WMRs) have been proven to be quite competitive and useful in outdoor missions. However, they may face serious sinkage or slippage on deformable terrains, and even get stuck or damaged, thereby causing mission failure. To mitigate these risks, it is essential to closely monitor

Cited by 10SourceScholar
2022

An Efficient Training Approach for Very Large Scale Face Recognition

CVPR 2022poster

Face recognition has achieved significant progress in deep learning era due to the ultra-large-scale and welllabeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is indispensable. Th…

Cited by 39PDFcodeScholar
2022

CAFE: Learning To Condense Dataset by Aligning Features

CVPR 2022poster

Dataset condensation aims at reducing the network training effort through condensing a cumbersome training set into a compact synthetic one. State-of-the-art approaches largely rely on learning the synthetic data by matching the gradients between the real and synthetic data batches. Despite the intu…

Cited by 277PDFcodeScholar
2022

Crafting Better Contrastive Views for Siamese Representation Learning

CVPR 2022oral

Recent self-supervised contrastive learning methods greatly benefit from the Siamese structure that aims at minimizing distances between positive pairs. For high performance Siamese representation learning, one of the keys is to design good contrastive pairs. Most previous works simply apply random…

Cited by 140PDFcodeScholar
2022

Decoupled Multi-Task Learning With Cyclical Self-Regulation for Face Parsing

CVPR 2022poster

This paper probes intrinsic factors behind typical failure cases (e.g spatial inconsistency and boundary confusion) produced by the existing state-of-the-art method in face parsing. To tackle these problems, we propose a novel Decoupled Multi-task Learning with Cyclical Self-Regulation (DML-CSR) for…

Cited by 43PDFcodeScholar
2022

DenseCLIP: Language-Guided Dense Prediction With Context-Aware Prompting

CVPR 2022poster

Recent progress has shown that large-scale pre-training using contrastive image-text pairs can be a promising alternative for high-quality visual representation learning from natural language supervision. Benefiting from a broader source of supervision, this new paradigm exhibits impressive transfer…

Cited by 678PDFcodeScholar
2022

Dimension Embeddings for Monocular 3D Object Detection

CVPR 2022poster

Most existing deep learning-based approaches for monocular 3D object detection directly regress the dimensions of objects and overlook their importance in solving the ill-posed problem. In this paper, we propose a general method to learn appropriate embeddings for dimension estimation in monocular 3…

Cited by 20PDFScholar
2022

Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis

ECCV 2022poster

"Talking head synthesis is an emerging technology with wide applications in film dubbing, virtual avatars and online education. Recent NeRF-based methods generate more natural talking videos, as they better capture the 3D structural information of faces. However, a specific model needs to be trained…

2022

MVSTER: Epipolar Transformer for Efficient Multi-View Stereo

ECCV 2022poster

"Learning-based Multi-View Stereo (MVS) methods warp source images into the reference camera frustum to form 3D volumes, which are fused as a cost volume to be regularized by subsequent networks. The fusing step plays a vital role in bridging 2D semantics and 3D spatial associations. However, previo…

2022

OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression

NeurIPS 2022accept

This paper presents a language-powered paradigm for ordinal regression. Existing methods usually treat each rank as a category and employ a set of weights to learn these concepts. These methods are easy to overfit and usually attain unsatisfactory performance as the learned concepts are mainly deriv…

2022

Predict the Rover Mobility Over Soft Terrain Using Articulated Wheeled Bevameter

RA-L 2022

Robot mobility is critical for mission success, especially in soft or deformable terrains, where the complex wheel-soil interaction mechanics often leads to excessive wheel slip and sinkage, causing the eventual mission failure. To improve the rover performance, online mobility prediction using visi

Cited by 21SourceScholar
2022

Shapley-NAS: Discovering Operation Contribution for Neural Architecture Search

CVPR 2022poster

In this paper, we propose a Shapley value based method to evaluate operation contribution (Shapley-NAS) for neural architecture search. Differentiable architecture search (DARTS) acquires the optimal architectures by optimizing the architecture parameters with gradient descent, which significantly r…

Cited by 57PDFcodeScholar
2022

SurroundDepth: Entangling Surrounding Views for Self-Supervised Multi-Camera Depth Estimation

CoRL 2022poster

Depth estimation from images serves as the fundamental step of 3D perception for autonomous driving and is an economical alternative to expensive depth sensors like LiDAR. The temporal photometric consistency enables self-supervised depth estimation without labels, further facilitating its applicati…

Cited by 87SourcecodeScholar
2021

Global Filter Networks for Image Classification

NeurIPS 2021poster

Recent advances in self-attention and pure multi-layer perceptrons (MLP) models for vision have shown great potential in achieving promising performance with fewer inductive biases. These models are generally based on learning interaction among spatial locations from raw data. The complexity of self…

2021

SIMPLE: SIngle-network with Mimicking and Point Learning for Bottom-up Human Pose Estimation

AAAI 2021technical

The practical application requests both accuracy and efficiency on multi-person pose estimation algorithms. But the high accuracy and fast inference speed are dominated by top-down methods and bottom-up methods respectively. To make a better trade-off between accuracy and efficiency, we propose a no…

Cited by 16SourcePDFScholar
2021

Structure-Aware Face Clustering on a Large-Scale Graph With 107 Nodes

CVPR 2021poster

Face clustering is a promising method for annotating unlabeled face images. Recent supervised approaches have boosted the face clustering accuracy greatly, however their performance is still far from satisfactory. These methods can be roughly divided into global-based and local-based ones. Global-ba…

Cited by 47PDFcodeScholar
2021

WebFace260M: A Benchmark Unveiling the Power of Million-Scale Deep Face Recognition

CVPR 2021poster

In this paper, we contribute a new million-scale face benchmark containing noisy 4M identities/260M faces (WebFace260M) and cleaned 2M identities/42M faces (WebFace42M) training data, as well as an elaborately designed time-constrained evaluation protocol. Firstly, we collect 4M name list and downlo…

Cited by 313PDFScholar
2020

The Devil Is in the Details: Delving Into Unbiased Data Processing for Human Pose Estimation

CVPR 2020poster

Recently, the leading performance of human pose estimation is dominated by top-down methods. Being a fundamental component in training and inference, data processing has not been systematically considered in pose estimation community, to the best of our knowledge. In this paper, we focus on this pro…

Cited by 288PDFcodeScholar
2019

Attention-Guided Unified Network for Panoptic Segmentation

CVPR 2019poster

This paper studies panoptic segmentation, a recently proposed task which segments foreground (FG) objects at the instance level as well as background (BG) contents at the semantic level. Existing methods mostly dealt with these two problems separately, but in this paper, we reveal the underlying rel…

Cited by 353PDFScholar
2018

Distractor-aware Siamese Networks for Visual Object Tracking

ECCV 2018poster

Recently, Siamese networks have drawn great attention in visual tracking community because of their balanced accuracy and speed. However, features used in most Siamese tracking approaches can only discriminate foreground from the non-semantic backgrounds. The semantic backgrounds are always consider…

2018

End-to-End Flow Correlation Tracking With Spatial-Temporal Attention

CVPR 2018poster

Discriminative correlation filters (DCF) with deep convolutional features have achieved favorable performance in recent tracking benchmarks. However, most of existing DCF trackers only consider appearance features of current frame, and hardly benefit from motion and inter-frame information. The lack…

2018

High Performance Visual Tracking With Siamese Region Proposal Network

CVPR 2018poster

Visual object tracking has been a fundamental topic in recent years and many deep learning based trackers have achieved state-of-the-art performance on multiple benchmarks. However, most of these trackers can hardly get top performance with real-time speed. In this paper, we propose the Siamese regi…

Cited by 3306SourcePDFScholar