← Search

Jun-Yan Zhu

68 accepted papers

2026

DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothing

CVPR 2026

Video diffusion models have achieved remarkable generative performance, but their substantial computational and memory costs pose significant challenges for deployment, especially on consumer GPUs. As recent advances in attention optimization mitigate previous computational bottlenecks, linear layer

Cited by 0SourceScholar
2026

FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models

ICML 2026poster

Diffusion models have become a dominant paradigm for high-quality generative modeling, while post-training is essential for adapting them to diverse downstream applications. However, post-training of large diffusion models is still challenging due to the prohibitive memory footprints and slow traini…

Cited by 0SourceScholar
2026

Learning an Image Editing Model without Image Editing Pairs

ICLR 2026poster

Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Curren…

Cited by 0SourcecodeScholar
2026

MotionStream: Real-Time Video Generation with Interactive Motion Controls

ICLR 2026oral

Current motion-conditioned video generation methods suffer from prohibitive latency (minutes per video) and non-causal processing that prevents real-time interaction. We present MotionStream, enabling sub-second latency with up to 29 FPS streaming generation on a single GPU. Our approach begins by a…

Cited by 0SourcecodeScholar
2026

Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization

CVPR 2026

Visual concept personalization aims to transfer only specific image attributes, such as identity, expression, lighting, and style, into unseen contexts. However, existing methods rely on holistic embeddings from general-purpose image encoders, which entangle multiple visual factors and make it diffi

Cited by 0SourcecodeScholar
2026

PAT3D: Physics-Augmented Text-to-3D Scene Generation

ICLR 2026poster

We introduce PAT3D, the first physics-augmented text-to-3D scene generation framework that integrates vision–language models with physics-based simulation to produce physically plausible, simulation-ready, and intersection-free 3D scenes. Given a text prompt, PAT3D generates 3D objects, infers their…

Cited by 0SourcecodeScholar
2026

Scaling Group Inference for Diverse and High-Quality Generation

ICLR 2026poster

Generative models typically sample outputs independently, and recent inference-time guidance and scaling algorithms focus on improving the quality of individual samples. However, in real-world applications, users are often presented with a set of multiple images (e.g., 4-8) for each prompt, where in…

Cited by 0SourcecodeScholar
2025

Efficient Autoregressive Shape Generation via Octree-Based Adaptive Tokenization

ICCV 2025poster

Many 3D generative models rely on variational autoencoders (VAEs) to learn compact shape representations. However, existing methods encode all shapes into a fixed-size token, disregarding the inherent variations in scale and complexity across 3D data. This leads to inefficient latent representations…

Cited by 0SourcePDFScholar
2025

Fast Data Attribution for Text-to-Image Models

NeurIPS 2025poster

Data attribution for text-to-image models aims to identify the training images that most significantly influenced a generated output. Existing attribution methods involve considerable computational resources for each query, making them impractical for real-world applications. We propose a novel app…

Cited by 0SourceScholar
2025

Generating Multi-Image Synthetic Data for Text-to-Image Customization

ICCV 2025poster

Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive test-time optimization or train encoders on single-image datasets without multi-image supervision, which can limit image…

2025

Generating Physically Stable and Buildable Brick Structures from Text

ICCV 2025poster

We introduce BrickGPT, the first approach for generating physically stable interconnecting brick assembly models from text prompts. To achieve this, we construct a large-scale, physically stable dataset of brick structures, along with their associated captions, and train an autoregressive large lang…

2025

Multi-subject Open-set Personalization in Video Generation

CVPR 2025poster

Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, require time-consuming optimization per subject, or support only a single subject. We present Video Alchemist--a video model…

Cited by 0SourcePDFScholar
2025

SVDQuant: Absorbing Outliers by Low-Rank Component for 4-Bit Diffusion Models

ICLR 2025spotlight

Diffusion models can effectively generate high-quality images. However, as they scale, rising memory demands and higher latency pose substantial deployment challenges. In this work, we aim to accelerate diffusion models by quantizing their weights and activations to 4 bits. At such an aggressive le…

2024

Co-speech Gesture Video Generation with 3D Human Meshes

ECCV 2024poster

"Co-speech gesture video generation is an enabling technique for many digital human applications. Substantial progress has been made in creating high-quality talking head videos. However, existing hand gesture video generation methods are primarily limited by the widely adopted 2D skeleton-based ges…

Cited by 1SourcePDFScholar
2024

CoFRIDA: Self-Supervised Fine-Tuning for Human-Robot Co-Painting

ICRA 2024poster

Prior robot painting and drawing work, such as FRIDA, has focused on decreasing the sim-to-real gap and expanding input modalities for users, but the interaction with these systems generally exists only in the input stages. To support interactive, human-robot collaborative painting, we introduce the…

Cited by 15SourcecodeScholar
2024

Data Attribution for Text-to-Image Models by Unlearning Synthesized Images

NeurIPS 2024poster

The goal of data attribution for text-to-image models is to identify the training images that most influence the generation of a new image. Influence is defined such that, for a given output, if a model is retrained from scratch without the most influential images, the model would fail to reproduce…

2024

Distilling Diffusion Models into Conditional GANs

ECCV 2024poster

"We propose a method to distill a complex multistep diffusion model into a single-step conditional GAN student model, dramatically accelerating inference, while preserving image quality. Our approach interprets diffusion distillation as a paired image-to-image translation task, using noise-to-image…

Cited by 39SourcePDFScholar
2024

FlashTex: Fast Relightable Mesh Texturing with LightControlNet

ECCV 2024oral

"Manually creating textures for 3D meshes is time-consuming, even for expert visual content creators. We propose a fast approach for automatically texturing an input 3D mesh based on a user-provided text prompt. Importantly, our approach disentangles lighting from surface material/reflectance in the…

Cited by 28SourcePDFScholar
2024

On the Content Bias in Frechet Video Distance

CVPR 2024poster

Frechet Video Distance (FVD) a prominent metric for evaluating video generation models is known to conflict with human perception occasionally. In this paper we aim to explore the extent of FVD's bias toward frame quality over temporal realism and identify its sources. We first quantify the FVD's se…

Cited by 35SourcePDFScholar
2024

Tactile DreamFusion: Exploiting Tactile Sensing for 3D Generation

NeurIPS 2024poster

3D generation methods have shown visually compelling results powered by diffusion image priors. However, they often fail to produce realistic geometric details, resulting in overly smooth surfaces or geometric details inaccurately baked in albedo maps. To address this, we introduce a new method that…

2023

Ablating Concepts in Text-to-Image Diffusion Models

ICCV 2023poster

Large-scale text-to-image diffusion models can generate high-fidelity images with powerful compositional ability. However, these models are typically trained on an enormous amount of Internet data, often containing copyrighted material, licensed images, and personal photos. Furthermore, they have be…

Cited by 212PDFcodeScholar
2023

Dense Text-to-Image Generation with Attention Modulation

ICCV 2023poster

Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image region. To address this, we propose DenseDiffusion, a training-free method that adapts a pre-trained text-to-image model t…

Cited by 125PDFcodeScholar
2023

Domain Expansion of Image Generators

CVPR 2023poster

Can one inject new concepts into an already trained generative model, while respecting its existing structure and knowledge? We propose a new task -- domain expansion -- to address this. Given a pretrained generator and novel (but related) domains, we expand the generator to jointly model all domain…

Cited by 17SourcePDFScholar
2023

Generalizing Dataset Distillation via Deep Generative Prior

CVPR 2023poster

Dataset Distillation aims to distill an entire dataset's knowledge into a few synthetic images. The idea is to synthesize a small number of synthetic data points that, when given to a learning algorithm as training data, result in a model approximating one trained on the original data. Despite a rec…

2023

Holistic Evaluation of Text-to-Image Models

NeurIPS 2023spotlight

The stunning qualitative improvement of text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (H…

2023

Multi-Concept Customization of Text-to-Image Diffusion

CVPR 2023poster

While generative models produce high-quality images of concepts learned from a large-scale database, a user often wishes to synthesize instantiations of their own concepts (for example, their family, pets, or items). Can we teach a model to quickly acquire a new concept, given a few examples? Furthe…

2023

Scaling Up GANs for Text-to-Image Synthesis

CVPR 2023highlight

The recent success of text-to-image synthesis has taken the world by storm and captured the general public's imagination. From a technical standpoint, it also marked a drastic change in the favored architecture to design generative image models. GANs used to be the de facto choice, with techniques l…

Cited by 613SourcePDFScholar
2023

Total-Recon: Deformable Scene Reconstruction for Embodied View Synthesis

ICCV 2023poster

We explore the task of embodied view synthesis from monocular videos of deformable scenes. Given a minute-long RGBD video of people interacting with their pets, we render the scene from novel camera trajectories derived from the in-scene motion of actors: (1) egocentric cameras that simulate the poi…

Cited by 20PDFcodeScholar
2022

Dataset Distillation by Matching Training Trajectories

CVPR 2022oral

Dataset distillation is the task of synthesizing a small dataset such that a model trained on the synthetic set will match the test accuracy of the model trained on the full dataset. The task is extremely challenging as it often involves backpropagating through the full training process or assuming…

Cited by 451PDFcodeScholar
2022

Depth-Supervised NeRF: Fewer Views and Faster Training for Free

CVPR 2022poster

A commonly observed failure mode of Neural Radiance Field (NeRF) is fitting incorrect geometries when given an insufficient number of input views. One potential reason is that standard volumetric rendering does not enforce the constraint that most of a scene's geometry consist of empty space and opa…

Cited by 1030PDFcodeScholar
2022

Efficient Spatially Sparse Inference for Conditional GANs and Diffusion Models

NeurIPS 2022accept

During image editing, existing deep generative models tend to re-synthesize the entire output from scratch, including the unedited regions. This leads to a significant waste of computation, especially for minor editing operations. In this work, we present Spatially Sparse Inference (SSI), a general-…

2022

GAN-Supervised Dense Visual Alignment

CVPR 2022oral

We propose GAN-Supervised Learning, a framework for learning discriminative models and their GAN-generated training data jointly end-to-end. We apply our framework to the dense visual alignment problem. Inspired by the classic Congealing method, our GANgealing algorithm trains a Spatial Transformer…

Cited by 78PDFcodeScholar
2022

SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

ICLR 2022poster

Guided image synthesis enables everyday users to create and edit photo-realistic images with minimum effort. The key challenge is balancing faithfulness to the user inputs (e.g., hand-drawn colored strokes) and realism of the synthesized images. Existing GAN-based methods attempt to achieve such bal…

2022

Spatially-Adaptive Multilayer Selection for GAN Inversion and Editing

CVPR 2022poster

Existing GAN inversion and editing methods work well for aligned objects with a clean background, such as portraits and animal faces, but often struggle for more difficult categories with complex scene layouts and object occlusions, such as cars, animals, and outdoor images. We propose a new method…

Cited by 49PDFcodeScholar
2021

Anycost GANs for Interactive Image Synthesis and Editing

CVPR 2021poster

Generative adversarial networks (GANs) have enabled photorealistic image synthesis and editing. However, due to the high computational cost of large-scale generators (e.g., StyleGAN2), it usually takes seconds to see the results of a single edit on edge devices, prohibiting interactive user experien…

Cited by 88PDFcodeScholar
2021

Editing Conditional Radiance Fields

ICCV 2021poster

A neural radiance field (NeRF) is a scene model supporting high-quality view synthesis, optimized per scene. In this paper, we explore enabling user editing of a category-level NeRF trained on a shape category. Specifically, we propose a method for propagating coarse 2D user scribbles to the 3D spac…

Cited by 300PDFcodeScholar
2020

Aligning and Projecting Images to Class-conditional Generative Networks

ECCV 2020poster

We present a method for projecting an input image into the space of a class-conditional generative neural network. We propose a method that optimizes for transformation to counteract the model biases in generative neural networks. Specifically, we demonstrate that one can solve for image translation…

Cited by 115SourcePDFScholar
2020

Differentiable Augmentation for Data-Efficient GAN Training

NeurIPS 2020poster

The performance of generative adversarial networks (GANs) heavily deteriorates given a limited amount of training data. This is mainly because the discriminatorsis memorizing the exact training set. To combat it, we propose Differentiable Augmentation (DiffAugment), a simple method that improves the…

2020

Diverse Image Generation via Self-Conditioned GANs

CVPR 2020poster

We introduce a simple but effective unsupervised method for generating diverse images. We train a class-conditional GAN model without using manually annotated class labels. Instead, our model is conditional on labels automatically derived from clustering in the discriminator's feature space. Our clu…

Cited by 132PDFcodeScholar
2020

GAN Compression: Efficient Architectures for Interactive Conditional GANs

CVPR 2020poster

Conditional Generative Adversarial Networks (cGANs) have enabled controllable image synthesis for many computer vision and graphics applications. However, recent cGANs are 1-2 orders of magnitude more computationally-intensive than modern recognition CNNs. For example, GauGAN consumes 281G MACs per…

Cited by 294PDFcodeScholar
2020

Swapping Autoencoder for Deep Image Manipulation

NeurIPS 2020poster

Deep generative models have become increasingly effective at producing realistic images from randomly sampled seeds, but using such models for controllable manipulation of existing images remains challenging. We propose the Swapping Autoencoder, a deep model designed specifically for image manipulat…

Cited by 403SourcePDFScholar
2020

The Hessian Penalty: A Weak Prior for Unsupervised Disentanglement

ECCV 2020poster

Existing popular methods for disentanglement rely on hand-picked priors and complex encoder-based architectures. In this paper, we propose the Hessian Penalty, a simple regularization function that encourages the input Hessian of a function to be diagonal. Our method is completely model-agnostic and…

2019

GAN Dissection: Visualizing and Understanding Generative Adversarial Networks

ICLR 2019poster

Generative Adversarial Networks (GANs) have recently achieved impressive results for many real-world applications, and many GAN variants have emerged with improvements in sample quality and training stability. However, visualization and understanding of GANs is largely missing. How does a GAN repres…

2019

Propagation Networks for Model-Based Control Under Partial Observation

ICRA 2019poster

There has been an increasing interest in learning dynamics simulators for model-based control. Compared with off-the-shelf physics engines, a learnable simulator can quickly adapt to unseen objects, scenes, and tasks. However, existing models like interaction networks only work for fully observable…

Cited by 170SourcecodeScholar
2019

Seeing What a GAN Cannot Generate

ICCV 2019oral

Despite the success of Generative Adversarial Networks (GANs), mode collapse remains a serious issue during GAN training. To date, little work has focused on understanding and quantifying which modes have been dropped by a model. In this work, we visualize mode collapse at both the distribution leve…

Cited by 457PDFcodeScholar
2019

Semantic Image Synthesis With Spatially-Adaptive Normalization

CVPR 2019oral

We propose spatially-adaptive normalization, a simple but effective layer for synthesizing photorealistic images given an input semantic layout. Previous methods directly feed the semantic layout as input to the network, forcing the network to memorize the information throughout all the layers. Inst…

Cited by 3706PDFcodeScholar
2018

3D-Aware Scene Manipulation via Inverse Graphics

NeurIPS 2018poster

We aim to obtain an interpretable, expressive, and disentangled scene representation that contains comprehensive structural and textural information for each object. Previous scene representations learned by neural networks are often uninterpretable, limited to a single object, or lacking 3D knowled…

2018

CyCADA: Cycle-Consistent Adversarial Domain Adaptation

ICML 2018oral

Domain adaptation is critical for success in new, unseen environments. Adversarial adaptation models have shown tremendous progress towards adapting to new environments by focusing either on discovering domain invariant representations or by mapping between unpaired image domains. While feature spac…

2018

High-Resolution Image Synthesis and Semantic Manipulation With Conditional GANs

CVPR 2018poster

We present a new method for synthesizing high-resolution photo-realistic images from semantic label maps using conditional generative adversarial networks (conditional GANs). Conditional GANs have enabled a variety of applications, but the results are often limited to low-resolution and still far fr…

2018

Video-to-Video Synthesis

NeurIPS 2018poster

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of the source video. While its image counterpart, the image-to…

2018

Visual Object Networks: Image Generation with Disentangled 3D Representations

NeurIPS 2018poster

Recent progress in deep generative models has led to tremendous breakthroughs in image generation. While being able to synthesize photorealistic images, existing models lack an understanding of our underlying 3D world. Different from previous works built on 2D datasets and models, we present a new g…

2017

Image-To-Image Translation With Conditional Adversarial Networks

CVPR 2017poster

We investigate conditional adversarial networks as a general-purpose solution to image-to-image translation problems. These networks not only learn the mapping from input image to output image, but also learn a loss function to train this mapping. This makes it possible to apply the same generic app…

Cited by 27229PDFcodeScholar
2017

Toward Multimodal Image-to-Image Translation

NeurIPS 2017poster

Many image-to-image translation problems are ambiguous, as a single input image may correspond to multiple possible outputs. In this work, we aim to model a distribution of possible outputs in a conditional generative modeling setting. The ambiguity of the mapping is distilled in a low-dimensional l…

2017

Unpaired Image-To-Image Translation Using Cycle-Consistent Adversarial Networks

ICCV 2017spotlight

Image-to-image translation is a class of vision and graphics problems where the goal is to learn the mapping between an input image and an output image using a training set of aligned image pairs. However, for many tasks, paired training data will not be available. We present an approach for learnin…

Cited by 27623PDFcodeScholar
2015

Learning a Discriminative Model for the Perception of Realism in Composite Images

ICCV 2015poster

What makes an image appear realistic? In this work, we are answering this question from a data-driven perspective by learning the perception of visual realism directly from large amounts of data. In particular, we train a Convolutional Neural Network (CNN) model that distinguishes natural photograph…

Cited by 174PDFcodeScholar