← Search

Chun-Hao P. Huang

15 accepted papers

2025

Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation

CVPR 2025poster

While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hypothesize that this is because there is no explicit supervision in terms of spat…

2025

VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors

CVPR 2025poster

Generative methods for image and video editing use generative models as priors to perform edits despite incomplete information, such as changing the composition of 3D objects shown in a single image. Recent methods have shown promising composition editing results in the image setting, but in the vid…

Cited by 1SourcePDFScholar
2023

3D Human Pose Estimation via Intuitive Physics

CVPR 2023poster

Estimating 3D humans from images often produces implausible bodies that lean, float, or penetrate the floor. Such methods ignore the fact that bodies are typically supported by the scene. A physics engine can be used to enforce physical plausibility, but these are not differentiable, rely on unreali…

Cited by 98SourcePDFScholar
2023

Blowing in the Wind: CycleNet for Human Cinemagraphs From Still Images

CVPR 2023poster

Cinemagraphs are short looping videos created by adding subtle motions to a static image. This kind of media is popular and engaging. However, automatic generation of cinemagraphs is an underexplored area and current solutions require tedious low-level manual authoring by artists. In this paper, we…

Cited by 16SourcePDFScholar
2023

MIME: Human-Aware 3D Scene Generation

CVPR 2023poster

Generating realistic 3D worlds occupied by moving humans has many applications in games, architecture, and synthetic data creation. But generating such scenes is expensive and labor intensive. Recent work generates human poses and motions given a 3D scene. Here, we take the opposite approach and gen…

2023

Reconstructing Signing Avatars From Video Using Linguistic Priors

CVPR 2023poster

Sign language (SL) is the primary method of communication for the 70 million Deaf people around the world. Video dictionaries of isolated signs are a core SL learning tool. Replacing these with 3D avatars can aid learning and enable AR/VR applications, improving access to technology and online media…

Cited by 15SourcePDFScholar
2023

SmartMocap: Joint Estimation of Human and Camera Motion Using Uncalibrated RGB Cameras

RA-L 2023

Markerless human motion capture (mocap) from multiple RGB cameras is a widely studied problem. Existing methods either need calibrated cameras or calibrate them relative to a static camera, which acts as the reference frame for the mocap system. The calibration step has to be done a priori for every

Cited by 13SourcecodeScholar
2022

Accurate 3D Body Shape Regression Using Metric and Semantic Attributes

CVPR 2022oral

While methods that regress 3D human meshes from images have progressed rapidly, the estimated body shapes often do not capture the true human shape. This is problematic since, for many applications, accurate body shape is as important as pose. The key reason that body shape accuracy lags pose accura…

Cited by 72PDFcodeScholar
2022

Capturing and Inferring Dense Full-Body Human-Scene Contact

CVPR 2022poster

Inferring human-scene contact (HSC) is the first step toward understanding how humans interact with their surroundings. While detecting 2D human-object interaction (HOI) and reconstructing 3D human pose and shape (HPS) have enjoyed significant progress, reasoning about 3D human-scene contact from a…

Cited by 150PDFScholar
2022

Human-Aware Object Placement for Visual Environment Reconstruction

CVPR 2022poster

Humans are in constant contact with the world as they move through it and interact with it. This contact is a vital source of information for understanding 3D humans, 3D scenes, and the interactions between them. In fact, we demonstrate that these human-scene interactions (HSIs) can be leveraged to…

Cited by 70PDFcodeScholar
2021

AGORA: Avatars in Geography Optimized for Regression Analysis

CVPR 2021poster

While the accuracy of 3D human pose estimation from images has steadily improved on benchmark datasets, the best methods still fail in many real-world scenarios. This suggests that there is a domain gap between current datasets and common scenes containing people. To obtain ground-truth 3D pose, cur…

Cited by 245PDFcodeScholar
2021

PARE: Part Attention Regressor for 3D Human Body Estimation

ICCV 2021poster

Despite significant progress, we show that state of the art 3D human pose and shape estimation methods remain sensitive to partial occlusion and can produce dramatically wrong predictions although much of the body is observable. To address this, we introduce a soft attention mechanism, called the Pa…

Cited by 480PDFcodeScholar
2021

SPEC: Seeing People in the Wild With an Estimated Camera

ICCV 2021poster

Due to the lack of camera parameter information for in-the-wild images, existing 3D human pose and shape (HPS) estimation methods make several simplifying assumptions: weak-perspective projection, large constant focal length, and zero camera rotation. These assumptions often do not hold and we show,…

Cited by 160PDFcodeScholar