← Search

Cheng Zhao

18 accepted papers

2026

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

ICML 2026poster

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality and semantic alignment of generated videos. However, recent…

Cited by 0SourceScholar
2026

Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding

CVPR 2026

Recent advances in 3D vision-language models (VLMs) highlight a strong potential for 3D scene understanding and reasoning.However, effectively tokenizing 3D scenes into holistic scene tokens, and leveraging these tokens across diverse 3D understanding tasks, remain highly challenging. We present NDT

Cited by 0SourcecodeScholar
2026

Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation

CVPR 2026

Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize that the view discrepancy arises from supervision limited to pixel- or latent-sp

Cited by 0SourceScholar
2025

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

CVPR 2025poster

Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challen…

2025

MMFN: Multi-Feature Multi-Modal Fusion Network for Diagnosis of Superficial Lymph Node Disease

ICASSP 2025accepted

The difficulty in identifying lymph node malignancies, including lymphoma and metastatic tumors, pose a diagnostic challenge at their primary sites. Given the heterogeneity of lymph node structures across different regions and the difficulty in distinguishing them from surrounding tissues, accurate…

Cited by 0SourceScholar
2025

SplatFlow: Self-Supervised Dynamic Gaussian Splatting in Neural Motion Flow Field for Autonomous Driving

CVPR 2025highlight

Most existing Dynamic Gaussian Splatting methods for complex dynamic urban scenarios rely on accurate object-level supervision from expensive manual labeling, limiting their scalability in real-world applications. In this paper, we introduce SplatFlow, a Self-Supervised Dynamic Gaussian Splatting wi…

Cited by 0SourcePDFScholar
2024

Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces Completion

CVPR 2024poster

In this paper we present a novel indoor 3D reconstruction method with occluded surface completion given a sequence of depth readings. Prior state-of-the-art (SOTA) methods only focus on the reconstruction of the visible areas in a scene neglecting the invisible areas due to the occlusions e.g. the c…

Cited by 1SourcePDFScholar
2024

SUP-NeRF: A Streamlined Unification of Pose Estimation and NeRF for Monocular 3D Object Reconstruction

ECCV 2024poster

"Monocular 3D reconstruction for categorical objects heavily relies on accurately perceiving each object’s pose. While gradient-based optimization in a NeRF framework updates the initial pose, this paper highlights that scale-depth ambiguity in monocular object reconstruction causes failures when th…

2023

3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D Detection

NeurIPS 2023poster

A major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets. While augmenting real scenes with virtual objects holds promise to improve both the diversity and quantity of the objects, it remains elusive due to the lack of an effective 3D objec…

2023

IF-Based Trajectory Planning and Cooperative Control for Transportation System of Cable Suspended Payload With Multi UAVs

IROS 2023poster

In this paper, we tackle the control and trajectory planning problems for the cooperative transportation system of cable-suspended payload with multi Unmanned Aerial Vehicles (UAVs). Firstly, a payload controller is presented considering the dynamic coupling between the UAV and the payload to accomp…

Cited by 3SourceScholar
2021

Monocular Teach-and-Repeat Navigation using a Deep Steering Network with Scale Estimation

IROS 2021poster

This paper proposes a novel monocular teach-and-repeat navigation system with the capability of scale awareness, i.e. the absolute distance between observation and goal images. It decomposes the navigation task into a sequence of visual servoing sub-tasks to approach consecutive goal/node images in…

Cited by 4SourceScholar
2021

NDT-Transformer: Large-Scale 3D Point Cloud Localisation using the Normal Distribution Transform Representation

ICRA 2021poster

3D point cloud-based place recognition is highly demanded by autonomous driving in GPS-challenged environments and serves as an essential component (i.e. loop-closure detection) in lidar-based SLAM systems. This paper proposes a novel approach, named NDT-Transformer, for real-time and large-scale pl…

Cited by 117SourcecodeScholar
2021

Robust and Long-term Monocular Teach and Repeat Navigation using a Single-experience Map

IROS 2021poster

This paper presents a robust monocular visual teach-and-repeat (VT&R) navigation system for long-term operation in outdoor environments. The approach leverages deep-learned descriptors to deal with the high illumination variance of the real world. In particular, a tailored self-supervised descriptor…

Cited by 11SourceScholar
2020

Generative Localization With Uncertainty Estimation Through Video-CT Data for Bronchoscopic Biopsy

RA-L 2020

Robot-assisted endobronchial intervention requires accurate localization based on both intra- and pre-operative data. Most existing methods achieve this by registering 2D videos with 3D CT models according to a defined similarity metric with local features. Instead, we formulate the bronchoscopic lo

Cited by 32SourceScholar
2019

Recurrent Kalman Networks: Factorized Inference in High-Dimensional Deep Feature Spaces

ICML 2019oral

In order to integrate uncertainty estimates into deep time-series modelling, Kalman Filters (KFs) (Kalman et al., 1960) have been integrated with deep learning models, however, such approaches typically rely on approximate inference tech- niques such as variational inference which makes learning mor…

2018

Learning Monocular Visual Odometry with Dense 3D Mapping from Dense 3D Flow

IROS 2018poster

This paper introduces a fully deep learning approach to monocular SLAM, which can perform simultaneous localization using a neural network for learning visual odometry (L-VO) and dense 3D mapping. Dense 2D flow and a depth image are generated from monocular images by sub-networks, which are then use…

Cited by 50SourceScholar
2018

Recurrent-OctoMap: Learning State-Based Map Refinement for Long-Term Semantic Mapping With 3-D-Lidar Data

RA-L 2018

This letter presents a novel semantic mapping approach, Recurrent-OctoMap, learned from long-term three-dimensional (3-D) Lidar data. Most existing semantic mapping approaches focus on improving semantic understanding of single frames, rather than 3-D refinement of semantic maps (i.e. fusing semanti

Cited by 77SourceScholar