← Search

Hyun Soo Park

28 accepted papers

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2023

Normal-Guided Garment UV Prediction for Human Re-Texturing

CVPR 2023highlight

Clothes undergo complex geometric deformations, which lead to appearance changes. To edit human videos in a physically plausible way, a texture map must take into account not only the garment transformation induced by the body movements and clothes fitting, but also its 3D fine-grained surface geome…

Cited by 15SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Learning Motion-Dependent Appearance for High-Fidelity Rendering of Dynamic Humans From a Single Camera

CVPR 2022poster

Appearance of dressed humans undergoes a complex geometric transformation induced not only by the static pose but also by its dynamics, i.e., there exists a number of cloth geometric configurations given a pose depending on the way it has moved. Such appearance modeling conditioned on motion has bee…

Cited by 17PDFScholar
2022

Learning To Detect Scene Landmarks for Camera Localization

CVPR 2022oral

Modern camera localization methods that use image retrieval, feature matching, and 3D structure-based pose estimation require long-term storage of numerous scene images or a vast amount of image features. This can make them unsuitable for resource constrained VR/AR devices and also raises serious pr…

Cited by 40PDFcodeScholar
2022

Look Both Ways: Self-Supervising Driver Gaze Estimation and Road Scene Saliency

ECCV 2022poster

"We present a new on-road driving dataset, called “Look Both Ways”, which contains synchronized video of both driver faces and the forward road scene, along with ground truth gaze data registered from eye tracking glasses worn by the drivers. Our dataset supports the study of methods for non-intrusi…

2022

Multiview Human Body Reconstruction from Uncalibrated Cameras

NeurIPS 2022accept

We present a new method to reconstruct 3D human body pose and shape by fusing visual features from multiview images captured by uncalibrated cameras. Existing multiview approaches often use spatial camera calibration (intrinsic and extrinsic parameters) to geometrically align and fuse visual feature…

Cited by 21SourcePDFScholar
2022

Self-supervised Wide Baseline Visual Servoing via 3D Equivariance

IROS 2022poster

One of the challenging input settings for visual servoing is when the initial and goal camera views are far apart. Such settings are difficult because the wide baseline can cause drastic changes in object appearance and cause occlusions. This paper presents a novel self-supervised visual servoing me…

Cited by 2SourceScholar
2021

Dense Keypoints via Multiview Supervision

NeurIPS 2021spotlight

This paper presents a new end-to-end semi-supervised framework to learn a dense keypoint detector using unlabeled multiview images. A key challenge lies in finding the exact correspondences between the dense keypoints in multiple views since the inverse of the keypoint mapping can be neither analytic…

Cited by 1SourcePDFScholar
2021

Inverse Simulation: Reconstructing Dynamic Geometry of Clothed Humans via Optimal Control

CVPR 2021poster

This paper studies the problem of inverse cloth simulation---to estimate shape and time-varying poses of the underlying body that generates physically plausible cloth motion, which matches to the point cloud measurements on the clothed humans. A key innovation is to represent the dynamics of the clo…

Cited by 16PDFScholar
2021

Pose-Guided Human Animation From a Single Image in the Wild

CVPR 2021poster

We present a new pose transfer method for synthesizing a human animation from a single image of a person controlled by a sequence of body poses. Existing pose transfer methods exhibit significant visual artifacts when applying to a novel scene, resulting in temporal inconsistency and failures in pre…

Cited by 77PDFScholar
2020

HUMBI: A Large Multiview Dataset of Human Body Expressions

CVPR 2020poster

This paper presents a new large multiview dataset called HUMBI for human body expressions with natural clothing. The goal of HUMBI is to facilitate modeling view-specific appearance and geometry of gaze, face, hand, body, and garment from assorted people. 107 synchronized HD cam- eras are used to ca…

Cited by 111PDFScholar
2020

Novel View Synthesis of Dynamic Scenes With Globally Coherent Depths From a Monocular Camera

CVPR 2020poster

This paper presents a new method to synthesize an image from arbitrary views and times given a collection of images of a dynamic scene. A key challenge for the novel view synthesis arises from dynamic scene reconstruction where epipolar geometry does not apply to the local motion of dynamic contents…

Cited by 172PDFScholar
2020

Surface Normal Estimation of Tilted Images via Spatial Rectifier

ECCV 2020poster

In this paper, we present a spatial rectifier to estimate surface normals of tilted images. Tilted images are of particular interest as more visual data are captured by arbitrarily oriented sensors such as body-/robot-mounted cameras. Existing approaches exhibit bounded performance on predicting sur…

2019

MONET: Multiview Semi-Supervised Keypoint Detection via Epipolar Divergence

ICCV 2019poster

This paper presents MONET---an end-to-end semi-supervised learning framework for a keypoint detector using multiview image streams. In particular, we consider general subjects such as non-human species where attaining a large scale annotated dataset is challenging. While multiview geometry can be us…

Cited by 63PDFcodeScholar
2019

Self-Supervised Adaptation of High-Fidelity Face Models for Monocular Performance Tracking

CVPR 2019oral

Improvements in data-capture and face modeling techniques have enabled us to create high-fidelity realistic face models. However, driving these realistic face models requires special input data, e.g., 3D meshes and unwrapped textures. Also, these face models expect clean input data taken under contr…

Cited by 43PDFScholar
2017

Am I a Baller? Basketball Performance Assessment From First-Person Videos

ICCV 2017poster

This paper presents a method to assess a basketball player's performance from his/her first-person video. A key challenge lies in the fact that the evaluation metric is highly subjective and specific to a particular evaluator. We leverage the first-person camera to address this challenge. The spatio…

Cited by 106PDFScholar
2017

Unsupervised Learning of Important Objects From First-Person Videos

ICCV 2017poster

A first-person camera, placed at a person's head, captures, which objects are important to the camera wearer. Most prior methods for this task learn to detect such important objects from the manually labeled first-person data in a supervised fashion. However, important objects are strongly related t…

Cited by 33PDFScholar