← Search

Zhiying Leng

5 accepted papers

2026

Cross-temporal 3D Gaussian Splatting for Sparse-view Guided Scene Update

AAAI 2026technical

Maintaining consistent 3D scene representations over time is a significant challenge in computer vision. Updating 3D scenes from sparse-view observations is crucial for various real-world applications, including urban planning, disaster assessment, and historical site preservation, where dense scan

Cited by 0SourcePDFScholar
2024

D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction

ECCV 2024poster

"Reconstructing hand-held objects from a single RGB image is a challenging task in computer vision. In contrast to prior works that utilize deterministic modeling paradigms, we employ a point cloud denoising diffusion model to account for the probabilistic nature of this problem. In the core, we int…

Cited by 2SourcePDFScholar
2024

HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape Generation

CVPR 2024poster

3D shape generation from text is a fundamental task in 3D representation learning. The text-shape pairs exhibit a hierarchical structure where a general text like "chair" covers all 3D shapes of the chair while more detailed prompts refer to more specific shapes. Furthermore both text and 3D shapes…

2023

Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction

ICCV 2023poster

Reconstructing both objects and hands in 3D from a single RGB image is complex. Existing methods rely on manually defined hand-object constraints in Euclidean space, leading to suboptimal feature learning. Compared with Euclidean space, hyperbolic space better preserves the geometric properties of m…

Cited by 14PDFScholar
2023

Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model

ICCV 2023poster

Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal and spatial relationships required to conform to a given text…

Cited by 53PDFScholar