← Search

Jae Shin Yoon

18 accepted papers

2025

Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

ICCV 2025poster

3D generation has made significant progress, however, it still largely remains at the object-level. Feedforward 3D scene-level generation has been rarely explored due to the lack of models capable of scaling-up latent representation learning on 3D scene-level data. Unlike object-level generative mod…

2025

Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and Harmonization

CVPR 2025poster

This paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricti…

Cited by 0SourcePDFScholar
2025

Text2Relight: Creative Portrait Relighting with Text Guidance

AAAI 2025technical

We present a lighting-aware image editing pipeline that, given a portrait image and a text prompt, performs single image relighting. Our model modifies the lighting and color of both the foreground and background to align with the provided text description. The unbounded nature in creativeness of a…

Cited by 1SourcePDFScholar
2024

Bidirectional Temporal Diffusion Model for Temporally Consistent Human Animation

ICLR 2024poster

We introduce a method to generate temporally coherent human animation from a single image, a video, or a random noise. This problem has been formulated as modeling of an auto-regressive generation, i.e., to regress past frames to decode future frames. However, such unidirectional generation is highl…

Cited by 1SourcePDFScholar
2024

COMPOSE: Comprehensive Portrait Shadow Editing

ECCV 2024poster

"Existing portrait relighting methods struggle with precise control over facial shadows, particularly when faced with challenges such as handling hard shadows from directional light sources or adjusting shadows while remaining in harmony with existing lighting conditions. In many situations, complet…

Cited by 3SourcePDFScholar
2024

High-Fidelity Modeling of Generalizable Wrinkle Deformation

ECCV 2024poster

"This paper proposes a generalizable model to synthesize high-fidelity clothing wrinkle deformation in 3D by learning from real data. Given the complex deformation behaviors of real-world clothing, this task presents significant challenges, primarily due to the lack of accurate ground-truth data. Ob…

Cited by 0SourcePDFScholar
2024

Relightful Harmonization: Lighting-aware Portrait Background Replacement

CVPR 2024poster

Portrait harmonization aims to composite a subject into a new background adjusting its lighting and color to ensure harmony with the background scene. Existing harmonization techniques often only focus on adjusting the global color and brightness of the foreground and ignore crucial illumination cue…

Cited by 15SourcePDFScholar
2024

Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction

CVPR 2024poster

This paper introduces the first text-guided work for generating the sequence of hand-object interaction in 3D. The main challenge arises from the lack of labeled data where existing ground-truth datasets are nowhere near generalizable in interaction type and object category which inhibits the modeli…

2023

Complete 3D Human Reconstruction From a Single Incomplete Image

CVPR 2023poster

This paper presents a method to reconstruct a complete human geometry and texture from an image of a person with only partial body observed, e.g., a torso. The core challenge arises from the occlusion: there exists no pixel to reconstruct where many existing single-view human reconstruction methods…

Cited by 16SourcePDFScholar
2022

Learning Motion-Dependent Appearance for High-Fidelity Rendering of Dynamic Humans From a Single Camera

CVPR 2022poster

Appearance of dressed humans undergoes a complex geometric transformation induced not only by the static pose but also by its dynamics, i.e., there exists a number of cloth geometric configurations given a pose depending on the way it has moved. Such appearance modeling conditioned on motion has bee…

Cited by 17PDFScholar
2021

Pose-Guided Human Animation From a Single Image in the Wild

CVPR 2021poster

We present a new pose transfer method for synthesizing a human animation from a single image of a person controlled by a sequence of body poses. Existing pose transfer methods exhibit significant visual artifacts when applying to a novel scene, resulting in temporal inconsistency and failures in pre…

Cited by 77PDFScholar
2020

HUMBI: A Large Multiview Dataset of Human Body Expressions

CVPR 2020poster

This paper presents a new large multiview dataset called HUMBI for human body expressions with natural clothing. The goal of HUMBI is to facilitate modeling view-specific appearance and geometry of gaze, face, hand, body, and garment from assorted people. 107 synchronized HD cam- eras are used to ca…

Cited by 111PDFScholar
2020

Novel View Synthesis of Dynamic Scenes With Globally Coherent Depths From a Monocular Camera

CVPR 2020poster

This paper presents a new method to synthesize an image from arbitrary views and times given a collection of images of a dynamic scene. A key challenge for the novel view synthesis arises from dynamic scene reconstruction where epipolar geometry does not apply to the local motion of dynamic contents…

Cited by 172PDFScholar
2019

Self-Supervised Adaptation of High-Fidelity Face Models for Monocular Performance Tracking

CVPR 2019oral

Improvements in data-capture and face modeling techniques have enabled us to create high-fidelity realistic face models. However, driving these realistic face models requires special input data, e.g., 3D meshes and unwrapped textures. Also, these face models expect clean input data taken under contr…

Cited by 43PDFScholar
2017

Pixel-Level Matching for Video Object Segmentation Using Convolutional Neural Networks

ICCV 2017poster

We propose a novel video object segmentation algorithm based on pixel-level matching using Convolutional Neural Networks (CNN). Our network aims to distinguish the target area from the background on the basis of the pixel-level similarity between two object units. The proposed network represents a t…

Cited by 219PDFScholar
2017

VPGNet: Vanishing Point Guided Network for Lane and Road Marking Detection and Recognition

ICCV 2017poster

In this paper, we propose a unified end-to-end trainable multi-task network that jointly handles lane and road marking detection and recognition that is guided by a vanishing point under adverse weather conditions. We tackle rainy and low illumination conditions, which have not been extensively stud…

Cited by 556PDFcodeScholar