← Search

Ruizhi Shao

16 accepted papers

2026

GAF: Gaussian Action Field As a 4D Representation for Dynamic World Modeling in Robotic Manipulation

ICRA 2026poster

Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action V-A paradigm, predicting actions directly from visual inputs, or a Vision-to-3D-to-Action V-3D-A paradigm, leveraging intermediate 3D representations. However, …

2025

ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping

CVPR 2025highlight

In this paper, we introduce ManiVideo, a novel method for generating consistent and temporally coherent bimanual hand-object manipulation videos from given motion sequences of hands and objects. The core idea of ManiVideo is the construction of a multi-layer occlusion (MLO) representation that learn…

Cited by 1SourcePDFScholar
2025

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios

NeurIPS 2025spotlight

Hand-Object Interaction (HOI) generation has significant application potential. However, current 3D HOI motion generation approaches heavily rely on predefined 3D object models and lab-captured motion data, limiting generalization capabilities. Meanwhile, HOI video generation methods prioritize pixe…

Cited by 0SourcecodeScholar
2025

The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

CVPR 2025poster

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for creating virtual characters that can communicate naturally in a…

Cited by 4SourcePDFScholar
2024

Control4D: Efficient 4D Portrait Editing with Text

CVPR 2024poster

We introduce Control4D an innovative framework for editing dynamic 4D portraits using text instructions. Our method addresses the prevalent challenges in 4D editing notably the inefficiencies of existing 4D representations and the inconsistent editing effect caused by diffusion-based editors. We fir…

Cited by 22SourcePDFScholar
2024

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior

ICLR 2024poster

We present DreamCraft3D, a hierarchical 3D content generation method that produces high-fidelity and coherent 3D objects. We tackle the problem by leveraging a 2D reference image to guide the stages of geometry sculpting and texture boosting. A central focus of this work is to address the consistenc…

2024

GPS-Gaussian: Generalizable Pixel-wise 3D Gaussian Splatting for Real-time Human Novel View Synthesis

CVPR 2024highlight

We present a new approach termed GPS-Gaussian for synthesizing novel views of a character in a real-time manner. The proposed method enables 2K-resolution rendering under a sparse-view camera setting. Unlike the original Gaussian Splatting or neural implicit rendering methods that necessitate per-su…

2024

HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion Models

CVPR 2024highlight

Recent years have witnessed a trend of the deep integration of the generation and reconstruction paradigms. In this paper we extend the ability of controllable generative models for a more comprehensive hand mesh recovery task: direct hand mesh generation inpainting reconstruction and fitting in a s…

Cited by 7SourcePDFScholar
2024

HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human Generation

CVPR 2024poster

Recent text-to-3D methods employing diffusion models have made significant advancements in 3D human generation. However these approaches face challenges due to the limitations of text-to-image diffusion models which lack an understanding of 3D structures. Consequently these methods struggle to achie…

2023

CloSET: Modeling Clothed Humans on Continuous Surface With Explicit Template Decomposition

CVPR 2023poster

Creating animatable avatars from static scans requires the modeling of clothing deformations in different poses. Existing learning-based methods typically add pose-dependent deformations upon a minimally-clothed mesh template or a learned implicit template, which have limitations in capturing detail…

Cited by 28SourcePDFScholar
2023

Tensor4D: Efficient Neural 4D Decomposition for High-Fidelity Dynamic Reconstruction and Rendering

CVPR 2023highlight

We present Tensor4D, an efficient yet effective approach to dynamic scene modeling. The key of our solution is an efficient 4D tensor decomposition method so that the dynamic scene can be directly represented as a 4D spatio-temporal tensor. To tackle the accompanying memory issue, we decompose the 4…

2022

DiffuStereo: High Quality Human Reconstruction via Diffusion-Based Stereo Using Sparse Cameras

ECCV 2022poster

"We propose DiffuStereo, a novel system using only sparse cameras (8 in this work) for high-quality 3D human reconstruction. At its core is a novel diffusion-based stereo module, which introduces diffusion models, a type of powerful generative models, into the iterative stereo matching network. To t…

Cited by 68SourcePDFScholar
2022

DoubleField: Bridging the Neural Surface and Radiance Fields for High-Fidelity Human Reconstruction and Rendering

CVPR 2022poster

We introduce DoubleField, a novel framework combining the merits of both surface field and radiance field for high-fidelity human reconstruction and rendering. Within DoubleField, the surface field and radiance field are associated together by a shared feature embedding and a surface-guided sampling…

Cited by 186PDFScholar
2022

Learning Implicit Templates for Point-Based Clothed Human Modeling

ECCV 2022poster

"We present FITE, a First-Implicit-Then-Explicit framework for modeling human avatars in clothing. Our framework first learns implicit surface templates representing the coarse clothing topology, and then employs the templates to guide the generation of point sets which further capture pose-dependen…

2021

DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview Cameras

ICCV 2021poster

We propose DeepMultiCap, a novel method for multi-person performance capture using sparse multi-view cameras. Our method can capture time varying surface details without the need of using pre-scanned template models. To tackle with the serious occlusion challenge for close interacting scenes, we com…

Cited by 110PDFScholar
2021

LocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography Estimation

ICCV 2021poster

Cross-resolution image alignment is a key problem in multiscale gigapixel photography, which requires to estimate homography matrix using images with large resolution gap. Existing deep homography methods concatenate the input images or features, neglecting the explicit formulation of correspondence…

Cited by 54PDFScholar