← Search

Zuozhuo Dai

8 accepted papers

2025

Tora: Trajectory-oriented Diffusion Transformer for Video Generation

CVPR 2025poster

Recent advancements in Diffusion Transformer (DiT) have demonstrated remarkable proficiency in producing high-quality video content. Nonetheless, the potential of transformer-based diffusion models for effectively generating videos with controllable motion remains an area of limited exploration. Thi…

2024

Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance

ECCV 2024poster

"In this study, we introduce a methodology for human image animation by leveraging a 3D human parametric model within a latent diffusion framework to enhance shape alignment and motion guidance in current human generative techniques. The methodology utilizes the SMPL(Skinned Multi-Person Linear) mod…

2024

Gaussian-Flow: 4D Reconstruction with Dynamic 3D Gaussian Particle

CVPR 2024highlight

We introduce Gaussian-Flow a novel point-based approach for fast dynamic scene reconstruction and real-time rendering from both multi-view and monocular videos. In contrast to the prevalent NeRF-based approaches hampered by slow training and rendering speeds our approach harnesses recent advancement…

Cited by 99SourcePDFScholar
2023

DRO: Deep Recurrent Optimizer for Video to Depth

RA-L 2023

There are increasing interests of studying the video-to-depth (V2D) problem with machine learning techniques. While earlier methods directly learn a mapping from images to depth maps and camera poses, more recent works enforce multi-view geometry constraints through optimization embedded in the lear

Cited by 21SourcecodeScholar
2022

Neural Window Fully-Connected CRFs for Monocular Depth Estimation

CVPR 2022poster

Estimating the accurate depth from a single image is challenging since it is inherently ambiguous and ill-posed. While recent works design increasingly complicated and powerful networks to directly regress the depth map, we take the path of CRFs optimization. Due to the expensive computation, CRFs a…

Cited by 424PDFScholar
2022

RCP: Recurrent Closest Point for Point Cloud

CVPR 2022oral

3D motion estimation including scene flow and point cloud registration has drawn increasing interest. Inspired by 2D flow estimation, recent methods employ deep neural networks to construct the cost volume for estimating accurate 3D flow. However, these methods are limited by the fact that it is dif…

Cited by 34PDFcodeScholar
2020

Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching

CVPR 2020oral

The deep multi-view stereo (MVS) and stereo matching approaches generally construct 3D cost volumes to regularize and regress the output depth or disparity. These methods are limited when high-resolution outputs are needed since the memory and time costs grow cubically as the volume resolution incre…

Cited by 897PDFcodeScholar
2019

Batch DropBlock Network for Person Re-Identification and Beyond

ICCV 2019poster

Since the person re-identification task often suffers from the problem of pose changes and occlusions, some attentive local features are often suppressed when training CNNs. In this paper, we propose the Batch DropBlock (BDB) Network which is a two branch network composed of a conventional ResNet-50…

Cited by 317PDFScholar