← Search

Minh Vo

15 accepted papers

2026

Clothe and Pose

CVPR 2026

We introduce Clothe and Pose, an image generation and editing task that enables users to try on garments while simultaneously adopting any desired pose. Our method takes a single user image, a set of garment images, and a reference pose as input, and outputs the user wearing the target garment in th

Cited by 0SourceScholar
2023

Ego-Humans: An Ego-Centric 3D Multi-Human Benchmark

ICCV 2023oral

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios, which limit the generalization of computer vision algorithms…

Cited by 39PDFScholar
2022

BANMo: Building Animatable 3D Neural Models From Many Casual Videos

CVPR 2022oral

Prior work for articulated 3D shape reconstruction often relies on specialized multi-view and depth sensors or pre-built deformable 3D models. Such methods do not scale to diverse sets of objects in the wild. We present a method that requires neither of them. It builds high-fidelity, articulated 3D…

Cited by 202PDFcodeScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

LISA: Learning Implicit Shape and Appearance of Hands

CVPR 2022poster

This paper proposes a do-it-all neural model of human hands, named LISA. The model can capture accurate hand shape and appearance, generalize to arbitrary hand subjects, provide dense surface correspondences, be reconstructed from images in the wild and easily animated. We train LISA by minimizing t…

Cited by 80PDFScholar
2022

TAVA: Template-Free Animatable Volumetric Actors

ECCV 2022poster

"Coordinate-based volumetric representations have the potential to generate photo-realistic virtual avatars from images. However, virtual avatars need to be controllable and be rendered in novel poses that may not have been observed. Traditional techniques, such as LBS, provide such a controlling fu…

2021

ANR: Articulated Neural Rendering for Virtual Avatars

CVPR 2021poster

Deferred Neural Rendering (DNR) uses a three-step pipeline to translate a mesh representation into an RGB image. The combination of a traditional rendering stack with neural networks hits a sweet spot in terms of computational complexity and realism of the resulting images. Using skinned meshes for…

Cited by 70PDFScholar
2021

ContactOpt: Optimizing Contact To Improve Grasps

CVPR 2021poster

Physical contact between hands and objects plays a critical role in human grasps. We show that optimizing the pose of a hand to achieve expected contact with an object can improve hand poses inferred via image-based methods. Given a hand mesh and an object mesh, a deep model trained on ground truth…

Cited by 147PDFcodeScholar
2021

ODAM: Object Detection, Association, and Mapping Using Posed RGB Video

ICCV 2021poster

Localizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system…

Cited by 34PDFcodeScholar
2020

4D Visualization of Dynamic Events From Unconstrained Multi-View Videos

CVPR 2020poster

We present a data-driven approach for 4D space-time visualization of dynamic events from videos captured by hand-held multiple cameras. Key to our approach is the use of self-supervised neural networks specific to the scene to compose static and dynamic aspects of an event. Though captured from disc…

Cited by 83PDFScholar
2020

Long-term Human Motion Prediction with Scene Context

ECCV 2020poster

Human movement is goal-directed and influenced by the spatial layout of the objects in the scene. To plan future human motion, it is crucial to perceive the environment -- imagine how hard it is to navigate a new room with lights off. Existing works on predicting human motion do not pay attention to…

2020

TexMesh: Reconstructing Detailed Human Texture and Geometry from RGB-D Video

ECCV 2020poster

We present TexMesh, a novel approach to reconstruct detailed human meshes with high-resolution full-body texture from RGB-D video. TexMesh enables high quality free-viewpoint rendering of humans. Given the RGB frames, the captured environment map, and the coarse per-frame human mesh from RGB-D track…

Cited by 53SourcePDFScholar
2019

Occlusion-Net: 2D/3D Occluded Keypoint Localization Using Graph Networks

CVPR 2019poster

We present Occlusion-Net, a framework to predict 2D and 3D locations of occluded keypoints for objects, in a largely self-supervised manner. We use an off-the-shelf detector as input (like MaskRCNN) that is trained only on visible key point annotations. This is the only supervision used in this work…

Cited by 92PDFScholar
2018

CarFusion: Combining Point Tracking and Part Detection for Dynamic 3D Reconstruction of Vehicles

CVPR 2018poster

Despite significant research in the area, reconstruction of multiple dynamic rigid objects (eg. vehicles) observed from wide-baseline, uncalibrated and unsynchronized cameras, remains hard. On one hand, feature tracking works well within each view but is hard to correspond across multiple cameras…

Cited by 100SourcePDFScholar