← Search

Yunzhi Zhang

25 accepted papers

2026

Coupled Diffusion Sampling for Training-Free Multi-View Image Editing

CVPR 2026

Given a collection of multi-view images, we perform consistent multi-view editing with a training-free framework using pre-trained 2D editing models and a generative multi-view model. While 2D editing models can independently edit each image in a set of multi-view images of a 3D scene, they do not m

Cited by 0SourceScholar
2026

Discovering Hybrid World Representations with Co-Evolving Foundation Models

AAAI 2026technical

This perspective article discusses an emerging research direction: to what extent can foundation models yield usable structure for modeling the physical world? We offer a Markovian formulation of structured world models and outline the notion of multi-level hybrid world representations that support

Cited by 0SourcePDFScholar
2025

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

CVPR 2025poster

Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images of a specific instance in novel contexts, i.e., "identity-preserving generation". This setting, along with many other ta…

Cited by 11SourcePDFScholar
2025

Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset

CVPR 2025highlight

We introduce Digital Twin Catalog (DTC), a new large-scale photorealistic 3D object digital twin dataset. A digital twin of a 3D object is a highly detailed, virtually indistinguishable representation of a physical object, accurately capturing its shape, appearance, physical properties, and other at…

2025

Model-Based Policy Adaptation for Closed-Loop End-to-end Autonomous Driving

NeurIPS 2025poster

End-to-end (E2E) autonomous driving models have demonstrated strong performance in open-loop evaluations but often suffer from cascading errors and poor generalization in closed-loop settings. To address this gap, we propose Model-based Policy Adaptation (MPA), a general framework that enhances the…

Cited by 0SourceScholar
2025

The Scene Language: Representing Scenes with Programs, Words, and Embeddings

CVPR 2025highlight

We introduce the Scene Language, a visual scene representation that concisely and precisely describes the structure, semantics, and identity of visual scenes. It represents a scene with three key components: a program that specifies the hierarchical and relational structure of entities in the scene,…

Cited by 8SourcePDFScholar
2025

Weakly-Supervised Learning of Dense Functional Correspondences

ICCV 2025poster

Establishing dense correspondences across image pairs is essential for tasks such as shape reconstruction and robot manipulation. In the challenging setting of matching across different categories, the function of an object, i.e., the effect that an object can cause on other objects, can guide how c…

Cited by 0SourcePDFScholar
2024

Learning the 3D Fauna of the Web

CVPR 2024poster

Learning 3D models of all animals in nature requires massively scaling up existing solutions. With this ultimate goal in mind we develop 3D-Fauna an approach that learns a pan-category deformable 3D animal model for more than 100 animal species jointly. One crucial bottleneck of modeling animals is…

Cited by 19SourcePDFScholar
2024

Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos

ECCV 2024poster

"We introduce a new method for learning a generative model of articulated 3D animal motions from raw, unlabeled online videos. Unlike existing approaches for 3D motion synthesis, our model requires no pose annotations or parametric shape models for training; it learns purely from a collection of unl…

Cited by 3SourcePDFScholar
2024

SHINOBI: Shape and Illumination using Neural Object Decomposition via BRDF Optimization In-the-wild

CVPR 2024poster

We present SHINOBI an end-to-end framework for the reconstruction of shape material and illumination from object images captured with varying lighting pose and background. Inverse rendering of an object based on unconstrained image collections is a long-standing challenge in computer vision and grap…

Cited by 6SourcePDFScholar
2024

ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image

CVPR 2024poster

We introduce a 3D-aware diffusion model ZeroNVS for single-image novel view synthesis for in-the-wild scenes. While existing methods are designed for single objects with masked backgrounds we propose new techniques to address challenges introduced by in-the-wild multi-object scenes with complex back…

2023

Holistic Evaluation of Text-to-Image Models

NeurIPS 2023spotlight

The stunning qualitative improvement of text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (H…

2023

MaskViT: Masked Visual Pre-Training for Video Prediction

ICLR 2023poster

The ability to predict future visual observations conditioned on past observations and motor commands can enable embodied agents to plan solutions to a variety of tasks in complex environments. This work shows that we can create good video prediction models by pre-training transformers via masked vi…

Cited by 135SourcePDFScholar
2023

Stanford-ORB: A Real-World 3D Object Inverse Rendering Benchmark

NeurIPS 2023poster

We introduce Stanford-ORB, a new real-world 3D Object inverse Rendering Benchmark. Recent advances in inverse rendering have enabled a wide range of real-world applications in 3D content generation, moving rapidly from research and commercial use cases to consumer devices. While the results continue…

2022

IKEA-Manual: Seeing Shape Assembly Step by Step

NeurIPS 2022accept

Human-designed visual manuals are crucial components in shape assembly activities. They provide step-by-step guidance on how we should move and connect different parts in a convenient and physically-realizable way. While there has been an ongoing effort in building agents that perform assembly tasks…

Cited by 19SourcePDFScholar
2022

Translating a Visual LEGO Manual to a Machine-Executable Plan

ECCV 2022poster

"We study the problem of translating an image-based, step-by-step assembly manual created by human designers into machine-interpretable instructions. We formulate this problem as a sequential prediction task: at each step, our model reads the manual, locates the components to be added to the current…

Cited by 22SourcePDFScholar