← Search

Jimei Yang

51 accepted papers

2025

Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and Harmonization

CVPR 2025poster

This paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricti…

Cited by 0SourcePDFScholar
2025

Rethinking Layered Graphic Design Generation with a Top-Down Approach

ICCV 2025poster

Graphic design is crucial for conveying ideas and messages. Designers usually organize their work into objects, backgrounds, and vectorized text layers to simplify editing. However, this workflow demands considerable expertise. With the rise of GenAI methods, an endless supply of high-quality graphi…

Cited by 0SourcePDFScholar
2024

ActAnywhere: Subject-Aware Video Background Generation

NeurIPS 2024poster

We study a novel problem to automatically generate video background that tailors to foreground subject motion. It is an important problem for the movie industry and visual effects community, which traditionally requires tedious manual efforts to solve. To this end, we propose ActAnywhere, a video di…

2024

Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models

ECCV 2024poster

"Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion customization, has not yet been well investigated. To address the c…

2024

Synergistic Global-space Camera and Human Reconstruction from Videos

CVPR 2024poster

Remarkable strides have been made in reconstructing static scenes or human bodies from monocular videos. Yet the two problems have largely been approached independently without much synergy. Most visual SLAM methods can only reconstruct camera trajectories and scene structures up to scale while most…

Cited by 3SourcePDFScholar
2024

Visual Layout Composer: Image-Vector Dual Diffusion Model for Design Layout Generation

CVPR 2024poster

This paper proposes an image-vector dual diffusion model for generative layout design. Distinct from prior efforts that mostly ignore element-level visual information our approach integrates the power of a pre-trained large image diffusion model to guide layout composition in a vector diffusion mode…

Cited by 5SourcePDFScholar
2023

ContactGen: Generative Contact Modeling for Grasp Generation

ICCV 2023poster

This paper presents a novel object-centric contact representation ContactGen for hand-object interaction. The ContactGen comprises 3 components: a contact map indicates the contact location, a part map represents the contact hand part, and a direction map tells the contact direction within each part…

Cited by 30PDFcodeScholar
2023

Normal-Guided Garment UV Prediction for Human Re-Texturing

CVPR 2023highlight

Clothes undergo complex geometric deformations, which lead to appearance changes. To edit human videos in a physically plausible way, a texture map must take into account not only the garment transformation induced by the body movements and clothes fitting, but also its 3D fine-grained surface geome…

Cited by 15SourcePDFScholar
2023

Putting People in Their Place: Affordance-Aware Human Insertion Into Scenes

CVPR 2023poster

We study the problem of inferring scene affordances by presenting a method for realistically inserting people into scenes. Given a scene image with a marked region and an image of a person, we insert the person into the scene while respecting the scene affordances. Our model can infer the set of rea…

2022

Audio-Driven Neural Gesture Reenactment With Video Motion Graphs

CVPR 2022poster

Human speech is often accompanied by body gestures including arm and hand gestures. We present a method that reenacts a high-quality video with gestures matching a target speech audio. The key idea of our method is to split and re-assemble clips from a reference video through a novel video motion gr…

Cited by 19PDFcodeScholar
2022

Learning Motion-Dependent Appearance for High-Fidelity Rendering of Dynamic Humans From a Single Camera

CVPR 2022poster

Appearance of dressed humans undergoes a complex geometric transformation induced not only by the static pose but also by its dynamics, i.e., there exists a number of cloth geometric configurations given a pose depending on the way it has moved. Such appearance modeling conditioned on motion has bee…

Cited by 17PDFScholar
2022

Learning Visibility for Robust Dense Human Body Estimation

ECCV 2022poster

"Estimating 3D human pose and shape from 2D images is a crucial yet challenging task. While prior methods with model-based representations can perform reasonably well on whole-body images, they often fail when parts of the body are occluded or outside the frame. Moreover, these results usually do no…

2022

Skeleton-Free Pose Transfer for Stylized 3D Characters

ECCV 2022poster

"We present the first method that automatically transfers poses between stylized 3D characters without skeletal rigging. In contrast to previous attempts to learn pose transformations on fixed or topology-equivalent skeleton templates, our method focuses on a novel scenario to handle skeleton-free c…

Cited by 41SourcePDFScholar
2021

HuMoR: 3D Human Motion Model for Robust Pose Estimation

ICCV 2021poster

We introduce HuMoR: a 3D Human Motion Model for Robust Estimation of temporal pose and shape. Though substantial progress has been made in estimating 3D human motion and shape from dynamic observations, recovering plausible pose sequences in the presence of noise and occlusions remains a challenge.…

Cited by 354PDFcodeScholar
2021

Learning to Sit: Synthesizing Human-Chair Interactions via Hierarchical Control

AAAI 2021technical

Recent progress on physics-based character animation has shown impressive breakthroughs on human motion synthesis, through imitating motion capture data via deep reinforcement learning. However, results have mostly been demonstrated on imitating a single distinct motion pattern, and do not generaliz…

Cited by 45SourcePDFScholar
2020

Contact and Human Dynamics from Monocular Video

ECCV 2020poster

Existing deep models predict 2D and 3D kinematic poses from video that are approximately accurate, but contain visible errors that violate physical constraints, such as feet penetrating the ground and bodies leaning at extreme angles. In this paper, we present a physics-based method for inferring 3D…

2020

High-Resolution Image Inpainting with Iterative Confidence Feedback and Guided Upsampling

ECCV 2020poster

Existing image inpainting methods often produce artifacts when dealing with large holes in real applications. To address this challenge, we propose an iterative inpainting method with a feedback mechanism. Specifically, we introduce a deep generative model which not only outputs an inpainting result…

Cited by 219SourcePDFScholar
2019

Free-Form Image Inpainting With Gated Convolution

ICCV 2019oral

We present a generative image inpainting system to complete images with free-form mask and guidance. The system is based on gated convolutions learned from millions of images without additional labelling efforts. The proposed gated convolution solves the issue of vanilla convolution that treats all…

Cited by 2386PDFcodeScholar
2019

FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB Images

ICCV 2019poster

Estimating 3D hand pose from single RGB images is a highly ambiguous problem that relies on an unbiased training dataset. In this paper, we analyze cross-dataset generalization when training on existing datasets. We find that approaches perform well on the datasets they are trained on, but do not ge…

Cited by 544PDFScholar
2019

LayoutGAN: Generating Graphic Layouts with Wireframe Discriminators

ICLR 2019poster

Layout is important for graphic design and scene generation. We propose a novel Generative Adversarial Network, called LayoutGAN, that synthesizes layouts by modeling geometric relations of different types of 2D elements. The generator of LayoutGAN takes as input a set of randomly-placed 2D graphic…

Cited by 262SourcePDFScholar
2019

On the Continuity of Rotation Representations in Neural Networks

CVPR 2019poster

In neural networks, it is often desirable to work with various representations of the same space. For example, 3D rotations can be represented with quaternions or Euler angles. In this paper, we advance a definition of a continuous representation, which can be helpful for training deep neural netwo…

Cited by 1552PDFScholar
2018

BodyNet: Volumetric Inference of 3D Human Body Shapes

ECCV 2018poster

Human shape estimation is an important task for video editing, animation and fashion industry. Predicting 3D human body shape from natural images, however, is highly challenging due to factors such as variation in human bodies, clothing and viewpoint. Prior methods addressing this problem typically…

Cited by 530SourcePDFScholar
2018

Flow-Grounded Spatial-Temporal Video Prediction from Still Images

ECCV 2018poster

Existing video prediction methods mainly rely on observing multiple historical frames or focus on predicting the next one-frame. In this work, we study the problem of generating consecutive multiple future frames by observing one single still image only. We formulate the multi-frame prediction task…

Cited by 159SourcePDFScholar
2018

Generative Image Inpainting With Contextual Attention

CVPR 2018poster

Recent deep learning based approaches have shown promising results for the challenging task of inpainting large missing regions in an image. These methods can generate visually plausible image structures and textures, but often create distorted structures or blurry textures inconsistent with surroun…

2018

MAttNet: Modular Attention Network for Referring Expression Comprehension

CVPR 2018poster

In this paper, we address referring expression comprehension: localizing an image region described by a natural language expression. While most recent work treats expressions as a single unit, we propose to decompose them into three modular components related to subject appearance, location, and re…

2018

Neural Kinematic Networks for Unsupervised Motion Retargetting

CVPR 2018poster

We propose a recurrent neural network architecture with a Forward Kinematics layer and cycle consistency based adversarial training objective for unsupervised motion retargetting. Our network captures the high-level properties of an input motion by the forward kinematics layer, and adapts them to a…

Cited by 261SourcePDFScholar
2018

PlaneNet: Piece-Wise Planar Reconstruction From a Single RGB Image

CVPR 2018poster

This paper proposes a deep neural network (DNN) for piece-wise planar depthmap reconstruction from a single RGB image. While DNNs have brought remarkable progress to single-image pixel-wise depth prediction, piece-wise planar depthmap reconstruction requires a structured geometry representation, an…

2017

3D-PRNN: Generating Shape Primitives With Recurrent Neural Networks

ICCV 2017poster

The success of various applications including robotics, digital content creation, and visualization demand a structured and abstract representation of the 3D world from limited sensor data. Inspired by the nature of human perception of 3D shapes as a collection of simple parts, we explore such an ab…

Cited by 224PDFcodeScholar
2017

Decomposing Motion and Content for Natural Video Sequence Prediction

ICLR 2017poster

We propose a deep neural network for the prediction of future frames in natural video sequences. To effectively handle complex evolution of pixels in videos, we propose to decompose the motion and content, two key components generating dynamics in videos. Our model is built upon the Encoder-Decoder…

Cited by 740SourceScholar
2017

Diversified Texture Synthesis With Feed-Forward Networks

CVPR 2017spotlight

Recent progresses on deep discriminative and generative modeling have shown promising results on texture synthesis. However, existing feed-forward based methods trade off generality for efficiency, which suffer from many issues, such as shortage of generality (i.e., build one network per texture), l…

Cited by 341PDFScholar
2017

Learning to Generate Long-term Future via Hierarchical Prediction

ICML 2017poster

We propose a hierarchical approach for making long-term predictions of future frames. To avoid inherent compounding errors in recursive pixel-level prediction, we propose to first estimate high-level structure in the input frames, then predict how that structure evolves in the future, and finally by…

Cited by 459SourcePDFScholar
2017

Material Editing Using a Physically Based Rendering Network

ICCV 2017spotlight

The ability to edit materials of objects in images is desirable by many content creators. However, this is an extremely challenging task as it requires to disentangle intrinsic physical properties of an image. We propose an end-to-end network architecture that replicates the forward image formation…

Cited by 105PDFScholar
2017

Predicting Scene Parsing and Motion Dynamics in the Future

NeurIPS 2017poster

It is important for intelligent systems, e.g. autonomous vehicles and robotics to anticipate the future in order to plan early and make decisions accordingly. Predicting the future scene parsing and motion dynamics helps the agents better understand the visual environment better as the former provid…

2017

Recurrent Multimodal Interaction for Referring Image Segmentation

ICCV 2017poster

In this paper we are interested in the problem of image segmentation given natural language descriptions, i.e. referring expressions. Existing works tackle this problem by first modeling images and sentences independently and then segment images by combining these two types of representations. We ar…

Cited by 296PDFcodeScholar
2017

Transformation-Grounded Image Generation Network for Novel 3D View Synthesis

CVPR 2017poster

We present a transformation-grounded image generation network for novel 3D view synthesis from a single image. Our approach first explicitly infers the parts of the geometry visible both in the input and novel views and then casts the remaining synthesis problem as image completion. Specifically, we…

Cited by 346PDFcodeScholar
2017

Universal Style Transfer via Feature Transforms

NeurIPS 2017poster

Universal style transfer aims to transfer arbitrary visual styles to content images. Existing feed-forward based methods, while enjoying the inference efficiency, are mainly limited by inability of generalizing to unseen styles or compromised visual quality. In this paper, we present a simple yet ef…

2017

Video Scene Parsing With Predictive Feature Learning

ICCV 2017poster

Video scene parsing is challenging due to the following two reasons: firstly, it is non-trivial to learn meaningful video representations for producing the temporally consistent labeling map; secondly, such a learning process becomes more difficult with insufficient labeled video training data. In t…

Cited by 156PDFScholar
2016

Object Contour Detection With a Fully Convolutional Encoder-Decoder Network

CVPR 2016spotlight

We develop a deep learning algorithm for contour detection with a fully convolutional encoder-decoder network. Different from previous low-level edge detection, our algorithm focuses on detecting higher-level object contours. Our network is trained end-to-end on PASCAL VOC with refined ground truth…

Cited by 478PDFScholar
2016

Object Tracking via Dual Linear Structured SVM and Explicit Feature Map

CVPR 2016poster

Structured support vector machine (SSVM) based methods has demonstrated encouraging performance in recent object tracking benchmarks. However, the complex and expensive optimization limits their deployment in real-world applications. In this paper, we present a simple yet efficient dual linear SSVM…

Cited by 297PDFScholar
2016

Perspective Transformer Nets: Learning Single-View 3D Object Reconstruction without 3D Supervision

NeurIPS 2016poster

Understanding the 3D world is a fundamental problem in computer vision. However, learning a good representation of 3D objects is still an open problem due to the high dimensionality of the data and many factors of variation involved. In this work, we investigate the task of single-view 3D object rec…

2015

PatchCut: Data-Driven Object Segmentation via Local Shape Transfer

CVPR 2015poster

Object segmentation is highly desirable for image understanding and editing. Current interactive tools require a great deal of user effort while automatic methods are usually limited to images of special object categories or with high color contrast. In this paper, we propose a data-driven algorithm…

Cited by 25SourcePDFScholar
2015

Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis

NeurIPS 2015poster

An important problem for both graphics and vision is to synthesize novel views of a 3D object from a single image. This is in particular challenging due to the partial observability inherent in projecting a 3D object onto the image space, and the ill-posedness of inferring object shape and pose. How…

Cited by 375SourcePDFScholar