← Search

Shangzhe Wu

23 accepted papers

2026

Particulate: Feed-Forward 3D Object Articulation

CVPR 2026

We introduce Particulate, a feed-forward model that, given a 3D mesh of an object, infers its articulations, including its 3D parts, their kinematic structure, and the motion constraints. The model is based on a transformer network, the Part Articulation Transformer, which predicts all these paramet

Cited by 0SourcecodeScholar
2025

DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction

CVPR 2025highlight

The choice of data representation is a key factor in the success of deep learning in geometric tasks. For instance, DUSt3R has recently introduced the concept of viewpoint- invariant point maps, generalizing depth prediction, and showing that one can reduce all the key problems in the 3D reconstruct…

Cited by 1SourcePDFScholar
2025

The Scene Language: Representing Scenes with Programs, Words, and Embeddings

CVPR 2025highlight

We introduce the Scene Language, a visual scene representation that concisely and precisely describes the structure, semantics, and identity of visual scenes. It represents a scene with three key components: a program that specifies the hierarchical and relational structure of entities in the scene,…

Cited by 8SourcePDFScholar
2024

Hearing Anything Anywhere

CVPR 2024poster

Recent years have seen immense progress in 3D computer vision and computer graphics with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However alongside immersive visual experiences immersive auditory experiences are equally vital to our…

2024

Learning the 3D Fauna of the Web

CVPR 2024poster

Learning 3D models of all animals in nature requires massively scaling up existing solutions. With this ultimate goal in mind we develop 3D-Fauna an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is…

Cited by 19SourcePDFScholar
2023

MagicPony: Learning Articulated 3D Animals in the Wild

CVPR 2023poster

We consider the problem of predicting the 3D shape, articulation, viewpoint, texture, and lighting of an articulated animal like a horse given a single test image as input. We present a new method, dubbed MagicPony, that learns this predictor purely from in-the-wild single-view images of the object…

2023

Stanford-ORB: A Real-World 3D Object Inverse Rendering Benchmark

NeurIPS 2023poster

We introduce Stanford-ORB, a new real-world 3D Object inverse Rendering Benchmark. Recent advances in inverse rendering have enabled a wide range of real-world applications in 3D content generation, moving rapidly from research and commercial use cases to consumer devices. While the results continue…

2022

Controllable 3D Face Synthesis with Conditional Generative Occupancy Fields

NeurIPS 2022accept

Capitalizing on the recent advances in image generation models, existing controllable face image synthesis methods are able to generate high-fidelity images with some levels of controllability, e.g., controlling the shapes, expressions, textures, and poses of the generated face images. However, thes…

Cited by 44SourcePDFScholar
2021

De-Rendering the World's Revolutionary Artefacts

CVPR 2021poster

Recent works have shown exciting results in unsupervised image de-rendering--learning to decompose 3D shape, appearance, and lighting from single-image collections without explicit supervision. However, many of these assume simplistic material and lighting models. We propose a method, termed RADAR,…

Cited by 35PDFcodeScholar
2021

Unsupervised Learning of Probably Symmetric Deformable 3D Objects from Images in the Wild (Extended Abstract)

IJCAI 2021poster

We propose a method to learn 3D deformable object categories from raw single-view images, without external supervision. The method is based on an autoencoder that factors each input image into depth, albedo, viewpoint and illumination. In order to disentangle these components without supervision, we…

2020

Self-Supervised Localisation between Range Sensors and Overhead Imagery

RSS 2020poster

Publicly available satellite imagery can be an ubiquitous, cheap, and powerful tool for vehicle localisation when a prior sensor map is unavailable. However, satellite images are not directly comparable to data from ground range sensors because of their starkly different modalities. We present a l…

Cited by 26SourcePDFScholar
2020

Unsupervised Learning of Probably Symmetric Deformable 3D Objects From Images in the Wild

CVPR 2020oral

We propose a method to learn 3D deformable object categories from raw single-view images, without external supervision. The method is based on an autoencoder that factors each input image into depth, albedo, viewpoint and illumination. In order to disentangle these components without supervision, we…

Cited by 367PDFcodeScholar
2018

Deep High Dynamic Range Imaging with Large Foreground Motions

ECCV 2018poster

This paper proposes the first non-flow-based deep framework for high dynamic range (HDR) imaging of dynamic scenes with large-scale foreground motions. In state-of-the-art deep HDR imaging, input images are first aligned using optical flows before merging, which are still error-prone due to occlusio…

2018

Image Generation from Sketch Constraint Using Contextual GAN

ECCV 2018poster

In this paper we investigate image generation guided by hand sketch. When the input sketch is badly drawn, the output of common image-to-image translation follows the input edges due to the hard condition imposed by the translation process. Instead, we propose to use sketch as weak constraint, where…

Cited by 174SourcePDFScholar