← Search

Chuanxia Zheng

30 accepted papers

2026

Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

CVPR 2026

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key contribution is a compact latent space that encodes the entire anim

Cited by 0SourcecodeScholar
2026

MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE

CVPR 2026

We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint representation of dense 3D point maps and 3D scene flows in a shared coordinate system, together with a 4D VAE tailored t

Cited by 0SourceScholar
2026

NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction

ICLR 2026poster

We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images, in a feed-forward manner. Unlike pixel-aligned methods that tie geometry to per-ray predictions, our formulation learns a global, view-agnostic scene representation that decouples reconstru…

Cited by 0SourcecodeScholar
2026

Particulate: Feed-Forward 3D Object Articulation

CVPR 2026

We introduce Particulate, a feed-forward model that, given a 3D mesh of an object, infers its articulations, including its 3D parts, their kinematic structure, and the motion constraints. The model is based on a transformer network, the Part Articulation Transformer, which predicts all these paramet

Cited by 0SourcecodeScholar
2025

Amodal3R: Amodal 3D Reconstruction from Occluded 2D Images

ICCV 2025poster

Most existing image-to-3D models assume that objects are fully visible, ignoring occlusions that commonly occur in real-world scenarios. In this paper, we introduce Amodal3R, a conditional image-to-3D model designed to reconstruct plausible 3D geometry and appearance from partial observations. We ex…

Cited by 0SourcePDFScholar
2025

DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness

ICCV 2025poster

Most 3D object generators prioritize aesthetic quality, often neglecting the physical constraints necessary for practical applications. One such constraint is that a 3D object should be self-supporting, i.e., remain balanced under gravity. Previous approaches to generating stable 3D objects relied o…

2025

Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

ICCV 2025poster

We introduce Geo4D, a method to repurpose video diffusion models for monocular 3D reconstruction of dynamic scenes. By leveraging the strong dynamic priors captured by large-scale pre-trained video models, Geo4D can be trained using only synthetic data while generalizing well to real data in a zero-…

2025

Puppet-Master: Scaling Interactive Video Generation as a Motion Prior for Part-Level Dynamics

ICCV 2025poster

We introduce Puppet-Master, an interactive video generator that captures the internal, part-level motion of objects, serving as a proxy for modeling object dynamics universally. Given an image of an object and a set of "drags" specifying the trajectory of a few points on the object, the model synthe…

Cited by 0SourcePDFScholar
2025

Semantix: An Energy-guided Sampler for Semantic Style Transfer

ICLR 2025poster

Recent advances in style and appearance transfer are impressive, but most methods isolate global style and local appearance transfer, neglecting semantic correspondence. Additionally, image and video tasks are typically handled in isolation, with little focus on integrating them for video transfer.…

Cited by 0SourcePDFScholar
2024

A General Protocol to Probe Large Vision Models for 3D Physical Understanding

NeurIPS 2024poster

Our objective in this paper is to probe large vision models to determine to what extent they ‘understand’ different physical properties of the 3D scene depicted in an image. To this end, we make the following contributions: (i) We introduce a general and lightweight protocol to evaluate whether feat…

2024

ClusteringSDF: Self-Organized Neural Implicit Surfaces for 3D Decomposition

ECCV 2024poster

"3D decomposition/segmentation remains a challenge as large-scale 3D annotated data is not readily available. Existing approaches typically leverage 2D machine-generated segments, integrating them to achieve 3D consistency. In this paper, we propose , a novel approach achieving both segmentation and…

Cited by 3SourcePDFScholar
2024

DragAPart: Learning a Part-Level Motion Prior for Articulated Objects

ECCV 2024poster

"We introduce , a method that, given an image and a set of drags as input, generates a new image of the same object that responds to the action of the drags. Differently from prior works that focused on repositioning objects, predicts part-level interactions, such as opening and closing a drawer. We…

Cited by 14SourcePDFScholar
2024

MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views

NeurIPS 2024poster

We introduce MVSplat360, a feed-forward approach for 360° novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and insufficient visual information provided, making it challenging for con…

2024

MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images

ECCV 2024oral

"We introduce , an efficient model that, given sparse multi-view images as input, predicts clean feed-forward 3D Gaussians. To accurately localize the Gaussian centers, we build a cost volume representation via plane sweeping, where the cross-view feature similarities stored in the cost volume can p…

2024

One More Step: A Versatile Plug-and-Play Module for Rectifying Diffusion Schedule Flaws and Enhancing Low-Frequency Controls

CVPR 2024poster

It is well known that many open-released foundational diffusion models have difficulty in generating images that substantially depart from average brightness despite such images being present in the training data. This is due to an inconsistency: while denoising starts from pure Gaussian noise durin…

Cited by 3SourcePDFScholar
2023

Cocktail: Mixing Multi-Modality Control for Text-Conditional Image Generation

NeurIPS 2023poster

Text-conditional diffusion models are able to generate high-fidelity images with diverse contents. However, linguistic representations frequently exhibit ambiguous descriptions of the envisioned objective imagery, requiring the incorporation of additional control signals to bolster the efficacy of t…

Cited by 23SourcePDFScholar
2023

Unified Discrete Diffusion for Simultaneous Vision-Language Generation

ICLR 2023poster

The recently developed discrete diffusion model performs extraordinarily well in generation tasks, especially in the text-to-image task, showing great potential for modeling multimodal signals. In this paper, we leverage these properties and present a unified multimodal generation model, which can p…

2023

Vector Quantized Wasserstein Auto-Encoder

ICML 2023poster

Learning deep discrete latent presentations offers a promise of better symbolic and summarized abstractions that are more useful to subsequent downstream tasks. Inspired by the seminal Vector Quantized Variational Auto-Encoder (VQ-VAE), most of work in learning deep discrete representations has main…

Cited by 18SourcePDFScholar
2022

Bridging Global Context Interactions for High-Fidelity Image Completion

CVPR 2022poster

Bridging global context interactions correctly is important for high-fidelity image completion with large masks. Previous methods attempting this via deep or large receptive field (RF) convolutions cannot escape from the dominance of nearby interactions, which may be inferior. In this paper, we prop…

Cited by 120PDFcodeScholar
2022

MoVQ: Modulating Quantized Vectors for High-Fidelity Image Generation

NeurIPS 2022accept

Although two-stage Vector Quantized (VQ) generative models allow for synthesizing high-fidelity and high-resolution images, their quantization operator encodes similar patches within an image into the same index, resulting in a repeated artifact for similar adjacent regions using existing decoder ar…

Cited by 86SourcePDFScholar
2022

Object-Compositional Neural Implicit Surfaces

ECCV 2022poster

"The neural implicit representation has shown its effectiveness in novel view synthesis and high-quality 3D reconstruction from multi-view images. However, most approaches focus on holistic scene representation yet ignore individual objects inside it, thus limiting potential downstream applications.…

2022

Sem2NeRF: Converting Single-View Semantic Masks to Neural Radiance Fields

ECCV 2022poster

"Image translation and manipulation have gain increasing attention along with the rapid development of deep generative models. Although existing approaches have brought impressive results, they mainly operated in 2D space. In light of recent advances in NeRF-based 3D-aware generative models, we intr…

2021

A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder

ICCV 2021poster

We present a unified and flexible framework to address the generalized problem of 3D motion synthesis that covers the tasks of motion prediction, completion, interpolation, and spatial-temporal recovery. Since these tasks have different input constraints and various fidelity and diversity requiremen…

Cited by 80PDFScholar
2018

T2Net: Synthetic-to-Realistic Translation for Solving Single-Image Depth Estimation Tasks

ECCV 2018poster

Current methods for single-image depth estimation use training datasets with real image-depth pairs or stereo pairs, which are not easy to acquire. We propose a framework, trained on synthetic image-depth pairs and unpaired real images, that comprises an image translation network for enhancing reali…