← Search

Yichun Shi

17 accepted papers

2026

Towards Robust Sequential Decomposition for Complex Image Editing

CVPR 2026

Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions. However, these models often struggle with complex instructions involving combinatorial editing operations or inter-step dependencies. This difficulty stems from the limitations of two c

Cited by 0SourceScholar
2026

VINCIE: Unlocking In-context Image Editing from Video

ICLR 2026poster

In-context image editing aims to modify images based on a contextual sequence comprising text and previously generated images. Existing methods typically depend on task-specific pipelines and expert models (e.g., segmentation and inpainting) to curate training data. In this work, we explore whether…

Cited by 0SourcecodeScholar
2025

Dual Diffusion for Unified Image Generation and Understanding

CVPR 2025poster

Diffusion models have gained tremendous success in text-to-image generation, yet still struggle with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end diffusion model for multi-modal understanding and generation that…

Cited by 81SourcePDFScholar
2025

HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

ICLR 2025poster

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely G…

Cited by 0SourcePDFScholar
2025

X-Dyna: Expressive Dynamic Human Image Animation

CVPR 2025highlight

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the subject and the surrounding environment. Building on prior appro…

2024

DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis

CVPR 2024highlight

We present DiffPortrait3D a conditional diffusion model that is capable of synthesizing 3D-consistent photo-realistic novel views from as few as a single in-the-wild portrait. Specifically given a single RGB input we aim to synthesize plausible but consistent facial details rendered from novel camer…

2024

Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences

CVPR 2024poster

Leveraging multi-view diffusion models as priors for 3D optimization have alleviated the problem of 3D consistency e.g. the Janus face problem or the content drift problem in zero-shot text-to-3D models. However the 3D geometric fidelity of the output remains an unresolved issue; albeit the rendered…

Cited by 1SourcePDFScholar
2024

MVDream: Multi-view Diffusion for 3D Generation

ICLR 2024poster

We introduce MVDream, a diffusion model that is able to generate consistent multi-view images from a given text prompt. Learning from both 2D and 3D data, a multi-view diffusion model can achieve the generalizability of 2D diffusion models and the consistency of 3D renderings. We demonstrate that su…

Cited by 630SourcePDFScholar
2024

MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion

ICML 2024poster

In this work, we propose MagicPose, a diffusion-based model for 2D human pose and facial expression retargeting. Specifically, given a reference image, we aim to generate a person's new images by controlling the poses and facial expressions while keeping the identity unchanged. To this end, we propo…

2023

OmniAvatar: Geometry-Guided Controllable 3D Head Synthesis

CVPR 2023poster

We present OmniAvatar, a novel geometry-guided 3D head synthesis model trained from in-the-wild unstructured images that is capable of synthesizing diverse identity-preserved 3D heads with compelling dynamic details under full disentangled control over camera poses, facial expressions, head shapes,…

Cited by 29SourcePDFScholar
2023

PAniC-3D: Stylized Single-View 3D Reconstruction From Portraits of Anime Characters

CVPR 2023poster

We propose PAniC-3D, a system to reconstruct stylized 3D character heads directly from illustrated (p)ortraits of (ani)me (c)haracters. Our anime-style domain poses unique challenges to single-view reconstruction; compared to natural images of human heads, character portrait illustrations have hair…

Cited by 21SourcePDFScholar
2023

PanoHead: Geometry-Aware 3D Full-Head Synthesis in 360deg

CVPR 2023poster

Synthesis and reconstruction of 3D human head has gained increasing interests in computer vision and computer graphics recently. Existing state-of-the-art 3D generative adversarial networks (GANs) for 3D human head synthesis are either limited to near-frontal views or hard to preserve 3D consistency…

2022

SemanticStyleGAN: Learning Compositional Generative Priors for Controllable Image Synthesis and Editing

CVPR 2022poster

Recent studies have shown that StyleGANs provide promising prior models for downstream tasks on image synthesis and editing. However, since the latent codes of StyleGANs are designed to control global styles, it is hard to achieve a fine-grained control over synthesized images. We present SemanticSt…

Cited by 111PDFcodeScholar
2020

Towards Universal Representation Learning for Deep Face Recognition

CVPR 2020poster

Recognizing wild faces is extremely hard as they appear with all kinds of variations. Traditional methods either train with specifically annotated variation data from target domains, or by introducing unlabeled target variation data to adapt from the training data. Instead, we propose a universal re…

Cited by 197PDFScholar