← Search

Hanbyul Joo

33 accepted papers

2026

Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer

ICLR 2026poster

We present Durian, the first method for generating portrait animation videos with cross-identity attribute transfer from one or more reference images to a target portrait. Training such models typically requires attribute pairs of the same individual, which are rarely available at scale. To address…

Cited by 0SourcecodeScholar
2026

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision

CVPR 2026

We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images, and a pose guidance video. Conventional two-stage pipelines treat image-based virtual try-on and pose-driven animation as separate processes, which oft

Cited by 0SourcecodeScholar
2025

DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

ICCV 2025poster

Modeling how humans interact with objects is crucial for AI to effectively assist or mimic human behaviors. Existing studies for learning such ability primarily focus on static human-object interaction (HOI) patterns, such as contact and spatial relationships, while dynamic HOI patterns, capturing t…

Cited by 0SourcePDFScholar
2025

HairCUP: Hair Compositional Universal Prior for 3D Gaussian Avatars

ICCV 2025poster

We present a universal prior model for 3D head avatars with explicit hair compositionality. Existing approaches to build generalizable priors for 3D head avatars often adopt a holistic modeling approach, treating the face and hair as an inseparable entity. This overlooks the inherent compositionalit…

Cited by 0SourcePDFScholar
2025

Learning 3D Object Spatial Relationships from Pre-trained 2D Diffusion Models

ICCV 2025poster

We present a method for learning 3D spatial relationships between object pairs, referred to as object-object spatial relationships (OOR), by leveraging synthetically generated 3D samples from pre-trained 2D diffusion models. We hypothesize that images synthesized by 2D diffusion models inherently ca…

Cited by 0SourcePDFScholar
2025

Learning to Generate Human-Human-Object Interactions from Textual Descriptions

NeurIPS 2025poster

The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines to understand such complex, context-dependent behaviors, it is essential to model multiple people in relation to the sur…

Cited by 0SourcecodeScholar
2025

ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions

CVPR 2025poster

To enable machines to understand the way humans interact with the physical world in daily life, 3D interaction signals should be captured in natural settings, allowing people to engage with multiple objects in a range of sequential and casual manipulations. To achieve this goal, we introduce our Par…

Cited by 17SourcePDFScholar
2024

GALA: Generating Animatable Layered Assets from a Single Scan

CVPR 2024poster

We present GALA a framework that takes as input a single-layer clothed 3D human mesh and decomposes it into complete multi-layered 3D assets. The outputs can then be combined with other assets to create novel clothed human avatars with any pose. Existing reconstruction approaches often treat clothed…

Cited by 7SourcePDFScholar
2024

Mocap Everyone Everywhere: Lightweight Motion Capture With Smartwatches and a Head-Mounted Camera

CVPR 2024poster

We present a lightweight and affordable motion capture method based on two smartwatches and a head-mounted camera. In contrast to the existing approaches that use six or more expert-level IMU devices our approach is much more cost-effective and convenient. Our method can make wearable motion capture…

Cited by 14SourcePDFScholar
2023

CHORUS : Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized Images

ICCV 2023oral

We present a method for teaching machines to understand and model the underlying spatial common sense of diverse human-object interactions in 3D in a self-supervised way. This is a challenging task, as there exist specific manifolds of the interactions that can be considered human-like and natural,…

Cited by 15PDFcodeScholar
2023

Chupa: Carving 3D Clothed Humans from Skinned Shape Priors using 2D Diffusion Probabilistic Models

ICCV 2023oral

We propose a 3D generation pipeline that uses diffusion models to generate realistic human digital avatars. Due to the wide variety of human identities, poses, and stochastic details, the generation of 3D human meshes has been a challenging problem. To address this, we decompose the problem into 2D…

Cited by 25PDFcodeScholar
2023

NCHO: Unsupervised Learning for Neural 3D Composition of Humans and Objects

ICCV 2023poster

Deep generative models have been recently extended to synthesizing 3D digital humans. However, previous approaches treat clothed humans as a single chunk of geometry without considering the compositionality of clothing and accessories. As a result, individual items cannot be naturally composed into…

Cited by 12PDFcodeScholar
2022

BANMo: Building Animatable 3D Neural Models From Many Casual Videos

CVPR 2022oral

Prior work for articulated 3D shape reconstruction often relies on specialized multi-view and depth sensors or pre-built deformable 3D models. Such methods do not scale to diverse sets of objects in the wild. We present a method that requires neither of them. It builds high-fidelity, articulated 3D…

Cited by 202PDFcodeScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Learning To Listen: Modeling Non-Deterministic Dyadic Facial Motion

CVPR 2022poster

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and speech audio of the speaker using a motion-audio cross attentio…

Cited by 105PDFScholar
2021

Body2Hands: Learning To Infer 3D Hands From Conversational Gesture Body Dynamics

CVPR 2021poster

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in non-verbal communication settings. We formulate the learning of th…

Cited by 55PDFcodeScholar
2020

3D Multi-bodies: Fitting Sets of Plausible 3D Human Models to Ambiguous Image Data

NeurIPS 2020spotlight

We consider the problem of obtaining dense 3D reconstructions of deformable objects from single and partially occluded views. In such cases, the visual evidence is usually insufficient to identify a 3D reconstruction uniquely, so we aim at recovering several plausible reconstructions compatible with…

Cited by 94SourcePDFScholar
2020

PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization

CVPR 2020oral

Recent advances in image-based 3D human shape estimation have been driven by the significant improvement in representation power afforded by deep neural networks. Although current approaches have demonstrated the potential in real world settings, they still fail to produce reconstructions with the l…

Cited by 955PDFcodeScholar
2020

Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

ECCV 2020poster

We present a method that infers spatial arrangements and shapes of humans and objects in a globally consistent 3D scene, all from a single image in-the-wild captured in an uncontrolled environment. Notably, our method runs on datasets without any scene- or object-level 3D supervision. Our key insigh…

2020

You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions

CVPR 2020oral

The body pose of a person wearing a camera is of great interest for applications in augmented reality, healthcare, and robotics, yet much of the person's body is out of view for a typical wearable camera. We propose a learning-based approach to estimate the camera wearer's 3D body pose from egocentr…

Cited by 113PDFcodeScholar
2019

Single-Network Whole-Body Pose Estimation

ICCV 2019poster

We present the first single-network approach for 2D whole-body pose estimation, which entails simultaneous localization of body, face, hands, and feet keypoints. Due to the bottom-up formulation, our method maintains constant real-time performance regardless of the number of people in the image. The…

Cited by 135PDFcodeScholar
2019

Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in a Triadic Interaction

CVPR 2019oral

We present a new research task and a dataset to understand human social interactions via computational methods, to ultimately endow machines with the ability to encode and decode a broad channel of social signals humans use. This research direction is essential to make a machine that genuinely commu…

Cited by 116PDFcodeScholar
2018

Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies

CVPR 2018poster

We present a unified deformation model for the markerless capture of multiple scales of human movement, including facial expressions, body motion, and hand gestures. An initial model is generated by locally stitching together models of the individual parts of the human body, which we refer to as the…

Cited by 617SourcePDFScholar
2017

Hand Keypoint Detection in Single Images Using Multiview Bootstrapping

CVPR 2017poster

We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedure multiview bootstrapping: first, an initial keypoint detector is used to produce noisy labels in multiple views of the…

Cited by 1579PDFScholar
2015

Panoptic Studio: A Massively Multiview System for Social Motion Capture

ICCV 2015oral

We present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social gr…

Cited by 1063PDFScholar