← Search

Xiaoguang Han

73 accepted papers

2026

EI-Part:Explode for Completion and Implode for Refinement

CVPR 2026

Part-level 3D generation is crucial for various downstream applications, including gaming, film production, and industrial design. However, decomposing a 3D shape into geometrically plausible and meaningful components remains a significant challenge. Previous part-based generation methods often stru

Cited by 0SourceScholar
2026

ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos

CVPR 2026

The ubiquity of monocular videos capturing daily hand-object interactions presents a valuable resource for embodied intelligence. While 3D hand reconstruction from in-the-wild videos has seen significant progress, reconstructing the involved objects remains challenging due to severe occlusions and t

Cited by 0SourcecodeScholar
2026

GarmentGPT: Compositional Garment Pattern Generation via Discrete Latent Tokenization

ICLR 2026poster

Apparel is a fundamental component of human appearance, making garment digitalization critical for digital human creation. However, sewing pattern creation traditionally relies on the intuition and extensive experience of skilled artisans. This manual bottleneck significantly hinders the scalability…

Cited by 0SourcecodeScholar
2026

LoFA: Learning to Predict Personalized Prior for Fast Adaptation of Visual Generative Models

CVPR 2026

Personalizing visual generative models to meet specific user needs has gained increasing attention, yet current methods like Low-Rank Adaptation (LoRA) remain impractical due to their demand for task-specific data and lengthy optimization. While a few hypernetwork-based approaches attempt to predict

Cited by 0SourcecodeScholar
2026

LumiTex: Towards High-Fidelity PBR Texture Generation with Illumination Context

ICLR 2026poster

Physically-based rendering (PBR) provides a principled standard for realistic material–lighting interactions in computer graphics. Despite recent advances in generating PBR textures, existing methods fail to address two fundamental challenges: 1) materials decomposition from image prompts under limi…

Cited by 0SourcecodeScholar
2026

MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE

CVPR 2026

We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint representation of dense 3D point maps and 3D scene flows in a shared coordinate system, together with a 4D VAE tailored t

Cited by 0SourceScholar
2026

OMGTex: One-stage Multi-style Facial Texture Reconstruction without Geometry Guidance

CVPR 2026

We propose OMGTex, an end-to-end diffusion-based framework for reconstructing high-quality and editable facial UV textures from multi-style facial images. Existing texture reconstruction methods face two major limitations: (1) Fragility due to reliance on 3D geometry priors, which are difficult to e

Cited by 0SourceScholar
2026

ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation

ICLR 2026poster

Existing multi-view 3D object reconstruction methods heavily rely on sufficient overlap between input views, where occlusions and sparse coverage in practice frequently yield severe reconstruction incompleteness. Recent advancements in diffusion-based 3D generative techniques offer the potential to…

Cited by 0SourcecodeScholar
2026

UniMo: Unified Motion Generation and Understanding with Chain of Thought

AAAI 2026technical

Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related tasks. While current unified frameworks based on large language models (LLMs) leverage linguistic priors, they frequently en

Cited by 0SourcePDFScholar
2026

UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents

CVPR 2026

Part-level 3D generation is essential for applications requiring decomposable and structured 3D synthesis. However, existing methods either rely on implicit part segmentation with limited granularity control or depend on strong external segmenters trained on large annotated datasets. In this work, w

Cited by 0SourceScholar
2025

Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging

ICCV 2025poster

With the growing demand for high-fidelity 3D models from 2D images, existing methods still face significant challenges in accurately reproducing fine-grained geometric details due to limitations in domain gaps and inherent ambiguities in RGB images. To address these issues, we propose Hi3DGen, a nov…

Cited by 0SourcePDFScholar
2025

HyPlaneHead: Rethinking Tri-plane-like Representations in Full-Head Image Synthesis

NeurIPS 2025poster

Tri-plane-like representations have been widely adopted in 3D-aware GANs for head image synthesis and other 3D object/scene modeling tasks due to their efficiency. However, querying features via Cartesian coordinate projection often leads to feature entanglement, which results in mirroring artifacts…

Cited by 0SourceScholar
2025

Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture

CVPR 2025poster

Existing methods for capturing multi-person holistic human motions from a monocular video usually involve integrating the detector, the tracker, and the human pose & shape estimator into a cascaded system. Differently, we develop a one-stage multi-person holistic human motion capture system, which 1…

2025

Robust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-On

CVPR 2025poster

This paper tackles the emerging challenge of multi-view virtual try-on, utilizing both front- and back-view clothing images as inputs. Extending frontal try-on methods to a multi-view context is not straightforward. Simply concatenating the two input views or encoding their features for a generative…

Cited by 0SourcePDFScholar
2025

Stable-SCore: A Stable Registration-based Framework for 3D Shape Correspondence

CVPR 2025poster

Establishing character shape correspondence is a critical and fundamental task in computer vision and graphics, with diverse applications including re-topology, attribute transfer, and shape interpolation. Current dominant functional map methods, while effective in controlled scenarios, struggle in…

Cited by 0SourcePDFScholar
2025

Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion

ICCV 2025poster

3D data simulation aims to bridge the gap between simulated and real-captured 3D data, which is a fundamental problem for real-world 3D visual tasks. Most 3D data simulation methods inject predefined physical priors but struggle to capture the full complexity of real data. An optimal approach involv…

Cited by 0SourcePDFScholar
2025

Towards Realistic Example-based Modeling via 3D Gaussian Stitching

CVPR 2025poster

Using parts of existing models to rebuild new models, commonly termed as example-based modeling, is a classical methodology in the realm of computer graphics. Previous works mostly focus on shape composition, making them very hard to use for realistic composition of 3D objects captured from real-wor…

Cited by 1SourcePDFScholar
2024

Free-ATM: Harnessing Free Attention Masks for Representation Learning on Diffusion-Generated Images

ECCV 2024poster

"This paper studies visual representation learning with diffusion-generated synthetic images. We start by uncovering that diffusion models’ cross-attention layers inherently provide annotation-free attention masks aligned with corresponding text inputs on generated images. We then investigate the pr…

2024

GaussReg: Fast 3D Registration with Gaussian Splatting

ECCV 2024poster

"Point cloud registration is a fundamental problem for large-scale 3D scene scanning and reconstruction. With the help of deep learning, registration methods have evolved significantly, reaching a nearly-mature stage. As the introduction of Neural Radiance Fields (NeRF), it has become the most popul…

Cited by 6SourcePDFScholar
2024

IPoD: Implicit Field Learning with Point Diffusion for Generalizable 3D Object Reconstruction from Single RGB-D Images

CVPR 2024highlight

Generalizable 3D object reconstruction from single-view RGB-D images remains a challenging task particularly with real-world data. Current state-of-the-art methods develop Transformer-based implicit field learning necessitating an intensive learning paradigm that requires dense query-supervision uni…

2024

LASA: Instance Reconstruction from Real Scans using A Large-scale Aligned Shape Annotation Dataset

CVPR 2024poster

Instance shape reconstruction from a 3D scene involves recovering the full geometries of multiple objects at the semantic instance level. Many methods leverage data-driven learning due to the intricacies of scene complexity and significant indoor occlusions. Training these methods often requires a l…

Cited by 4SourcePDFScholar
2024

Leveraging Noisy Labels of Nearest Neighbors for Label Correction and Sample Selection

ICASSP 2024accepted

Dealing with noisy labels (LNL) emerges as a critical challenge when applying deep learning (DL) in practical settings. Previous methodologies primarily concentrated on harnessing model predictions to mitigate the impact of noisy labels. Nevertheless, their efficacy is strongly contingent on the acc…

Cited by 0SourceScholar
2024

MVHumanNet: A Large-scale Dataset of Multi-view Daily Dressing Human Captures

CVPR 2024poster

In this era the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However in the realm of 3D vision while remarkable progress has been made with models trained on large-scale synthetic and real-captured object data like Objavers…

Cited by 19SourcePDFScholar
2024

PICTURE: PhotorealistIC virtual Try-on from UnconstRained dEsigns

CVPR 2024poster

In this paper we propose a novel virtual try-on from unconstrained designs (ucVTON) task to enable photorealistic synthesis of personalized composite clothing on input human images. Unlike prior arts constrained by specific input types our method allows flexible specification of style (text or image…

Cited by 8SourcePDFScholar
2024

RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D

CVPR 2024highlight

Lifting 2D diffusion for 3D generation is a challenging problem due to the lack of geometric prior and the complex entanglement of materials and lighting in natural images. Existing methods have shown promise by first creating the geometry through score-distillation sampling (SDS) applied to rendere…

2024

UPS: Unified Projection Sharing for Lightweight Single-Image Super-resolution and Beyond

NeurIPS 2024poster

To date, transformer-based frameworks have demonstrated impressive results in single-image super-resolution (SISR). However, under practical lightweight scenarios, the complex interaction of deep image feature extraction and similarity modeling limits the performance of these methods, since they req…

Cited by 1SourcePDFScholar
2024

Unveiling Advanced Frequency Disentanglement Paradigm for Low-Light Image Enhancement

ECCV 2024poster

"Previous low-light image enhancement (LLIE) approaches, while employing frequency decomposition techniques to address the intertwined challenges of low frequency (e.g., illumination recovery) and high frequency (e.g., noise reduction), primarily focused on the development of dedicated and complex n…

2023

A Comprehensive Benchmark for Neural Human Radiance Fields

NeurIPS 2023poster

The past two years have witnessed a significant increase in interest concerning NeRF-based human body rendering. While this surge has propelled considerable advancements, it has also led to an influx of methods and datasets. This explosion complicates experimental settings and makes fair comparisons…

Cited by 2SourcePDFScholar
2023

AUGUST: an Automatic Generation Understudy for Synthesizing Conversational Recommendation Datasets

ACL 2023findings

High-quality data is essential for conversational recommendation systems and serves as the cornerstone of the network architecture development and training strategy design. Existing works contribute heavy human efforts to manually labeling or designing and extending recommender dialogue templates. H…

2023

Activate and Reject: Towards Safe Domain Generalization under Category Shift

ICCV 2023poster

Albeit the notable performance on in-domain test points, it is non-trivial for deep neural networks to attain satisfactory accuracy when deploying in the open world, where novel domains and object classes often occur. In this paper, we study a practical problem of Domain Generalization under Categor…

Cited by 8PDFScholar
2023

CODA: Generalizing to Open and Unseen Domains with Compaction and Disambiguation

NeurIPS 2023spotlight

The generalization capability of machine learning systems degenerates notably when the test distribution drifts from the training distribution. Recently, Domain Generalization (DG) has been gaining momentum in enabling machine learning models to generalize to unseen domains. However, most DG methods…

Cited by 5SourcePDFScholar
2023

Efficient View Synthesis with Neural Radiance Distribution Field

ICCV 2023poster

Recent work on Neural Radiance Fields (NeRF) has demonstrated significant advances in high-quality view synthesis. A major limitation of NeRF is its low rendering efficiency due to the need for multiple network forwardings to render a single pixel. Existing methods to improve NeRF either reduce the…

Cited by 1PDFcodeScholar
2023

Exploring Motion Ambiguity and Alignment for High-Quality Video Frame Interpolation

CVPR 2023poster

For video frame interpolation(VFI), existing deep-learning-based approaches strongly rely on the ground-truth (GT) intermediate frames, which sometimes ignore the non-unique nature of motion judging from the given adjacent frames. As a result, these methods tend to produce averaged solutions that ar…

Cited by 26SourcePDFScholar
2023

Get3DHuman: Lifting StyleGAN-Human into a 3D Generative Model Using Pixel-Aligned Reconstruction Priors

ICCV 2023poster

Fast generation of high-quality 3D digital humans is important to a vast number of applications ranging from entertainment to professional concerns. Recent advances in differentiable rendering have enabled the training of 3D generative models without requiring 3D ground truths. However, the quality…

Cited by 24PDFScholar
2023

HairStep: Transfer Synthetic to Real Using Strand and Depth Maps for Single-View 3D Hair Modeling

CVPR 2023highlight

In this work, we tackle the challenging problem of learning-based single-view 3D hair modeling. Due to the great difficulty of collecting paired real image and 3D hair data, using synthetic data to provide prior knowledge for real domain becomes a leading solution. This unfortunately introduces the…

Cited by 26SourcePDFScholar
2023

MIMO Is All You Need:A Strong Multi-in-Multi-Out Baseline for Video Prediction

AAAI 2023technical

The mainstream of the existing approaches for video prediction builds up their models based on a Single-In-Single-Out (SISO) architecture, which takes the current frame as input to predict the next frame in a recursive manner. This way often leads to severe performance degradation when they try to e…

2023

MM-3DScene: 3D Scene Understanding by Customizing Masked Modeling With Informative-Preserved Reconstruction and Self-Distilled Consistency

CVPR 2023poster

Masked Modeling (MM) has demonstrated widespread success in various vision challenges, by reconstructing masked visual patches. Yet, applying MM for large-scale 3D scenes remains an open problem due to the data sparsity and scene complexity. The conventional random masking paradigm used in 2D images…

Cited by 12SourcePDFScholar
2023

MVImgNet: A Large-Scale Dataset of Multi-View Images

CVPR 2023poster

Being data-driven is one of the most iconic properties of deep learning algorithms. The birth of ImageNet drives a remarkable trend of "learning from large-scale data" in computer vision. Pretraining on ImageNet to obtain rich universal representations has been manifested to benefit various 2D visua…

Cited by 180SourcePDFScholar
2023

NeRFLix: High-Quality Neural View Synthesis by Learning a Degradation-Driven Inter-Viewpoint MiXer

CVPR 2023poster

Neural radiance fields(NeRF) show great success in novel-view synthesis. However, in real-world scenes, recovering high-quality details from the source images is still challenging for the existing NeRF-based approaches, due to the potential imperfect calibration information and scene representation…

2023

NerVE: Neural Volumetric Edges for Parametric Curve Extraction From Point Cloud

CVPR 2023poster

Extracting parametric edge curves from point clouds is a fundamental problem in 3D vision and geometry processing. Existing approaches mainly rely on keypoint detection, a challenging procedure that tends to generate noisy output, making the subsequent edge extraction error-prone. To address this is…

2023

REC-MV: REconstructing 3D Dynamic Cloth From Monocular Videos

CVPR 2023poster

Reconstructing dynamic 3D garment surfaces with open boundaries from monocular videos is an important problem as it provides a practical and low-cost solution for clothes digitization. Recent neural rendering methods achieve high-quality dynamic clothed human reconstruction results from monocular vi…

2023

RaBit: Parametric Modeling of 3D Biped Cartoon Characters With a Topological-Consistent Dataset

CVPR 2023poster

Assisting people in efficiently producing visually plausible 3D characters has always been a fundamental research topic in computer vision and computer graphics. Recent learning-based approaches have achieved unprecedented accuracy and efficiency in the area of 3D real human digitization. However, n…

Cited by 10SourcePDFScholar
2023

SCoDA: Domain Adaptive Shape Completion for Real Scans

CVPR 2023poster

3D shape completion from point clouds is a challenging task, especially from scans of real-world objects. Considering the paucity of 3D shape ground truths for real scans, existing works mainly focus on benchmarking this task on synthetic data, e.g. 3D computer-aided design models. However, the doma…

2022

Compound Domain Generalization via Meta-Knowledge Encoding

CVPR 2022poster

Domain generalization (DG) aims to improve the generalization performance for an unseen target domain by using the knowledge of multiple seen source domains. Mainstream DG methods typically assume that the domain label of each source sample is known a priori, which is challenged to be satisfied in m…

Cited by 84PDFScholar
2022

DArch: Dental Arch Prior-Assisted 3D Tooth Instance Segmentation With Weak Annotations

CVPR 2022poster

Automatic tooth instance segmentation on 3D dental models is a fundamental task for computer-aided orthodontic treatments. Existing learning-based methods rely heavily on expensive point-wise annotations. To alleviate this problem, we are the first to explore a low-cost annotation way for 3D tooth i…

Cited by 37PDFScholar
2022

ETHSeg: An Amodel Instance Segmentation Network and a Real-World Dataset for X-Ray Waste Inspection

CVPR 2022poster

Waste inspection for packaged waste is an important step in the pipeline of waste disposal. Previous methods either rely on manual visual checking or RGB image-based inspection algorithm, requiring costly preparation procedures (e.g., open the bag and spread the waste items). Moreover, occluded item…

Cited by 16PDFcodeScholar
2022

Expressive Talking Head Generation With Granular Audio-Visual Control

CVPR 2022poster

Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking H…

Cited by 148PDFScholar
2022

Multi-level Consistency Learning for Semi-supervised Domain Adaptation

IJCAI 2022poster

Semi-supervised domain adaptation (SSDA) aims to apply knowledge learned from a fully labeled source domain to a scarcely labeled target domain. In this paper, we propose a Multi-level Consistency Learning (MCL) framework for SSDA. Specifically, our MCL regularizes the consistency of different views…

2022

Pose2Room: Understanding 3D Scenes from Human Activities

ECCV 2022poster

"With wearable IMU sensors, one can estimate human poses from wearable devices without requiring visual input. In this work, we pose the question: Can we reason about object structure in real-world environments solely from human trajectory information? Crucially, we observe that human motion and int…

Cited by 17SourcePDFScholar
2022

Registering Explicit to Implicit: Towards High-Fidelity Garment Mesh Reconstruction From Single Images

CVPR 2022poster

Fueled by the power of deep learning techniques and implicit shape learning, recent advances in single-image human digitalization have reached unprecedented accuracy and could recover fine-grained surface details such as garment wrinkles. However, a common problem for the implicit-based methods is t…

Cited by 36PDFScholar
2022

SharpContour: A Contour-Based Boundary Refinement Approach for Efficient and Accurate Instance Segmentation

CVPR 2022poster

Excellent performance has been achieved on instance segmentation but the quality on the boundary area remains unsatisfactory, which leads to a rising attention on boundary refinement. For practical use, an ideal post-processing refinement scheme are required to be accurate, generic and efficient. Ho…

Cited by 34PDFScholar
2022

TO-Scene: A Large-Scale Dataset for Understanding 3D Tabletop Scenes

ECCV 2022poster

"Many basic indoor activities such as eating or writing are always conducted upon different tabletops (e.g., coffee tables, writing desks). It is indispensable to understanding tabletop scenes in 3D indoor scene parsing applications. Unfortunately, it is hard to meet this demand by directly deployin…

2022

Towards High-Fidelity Single-View Holistic Reconstruction of Indoor Scenes

ECCV 2022poster

"We present a new framework to reconstruct holistic 3D indoor scenes including both room background and indoor objects from single-view images. Existing methods can only produce 3D shapes of indoor objects with limited geometry quality because of the heavy occlusion of indoor scenes. To solve this,…

2021

3DCaricShop: A Dataset and a Baseline Method for Single-View 3D Caricature Face Reconstruction

CVPR 2021poster

Caricature is an artistic representation that deliberately exaggerates the distinctive features of a human face to convey humor or sarcasm. However, reconstructing a 3D caricature from a 2D caricature image remains a challenging task, mostly due to the lack of data. We propose to fill this gap by in…

Cited by 27PDFScholar
2021

LapsCore: Language-Guided Person Search via Color Reasoning

ICCV 2021poster

The key point of language-guided person search is to construct the cross-modal association between visual and textual input. Existing methods focus on designing multimodal attention mechanisms and novel cross-modal loss functions to learn such association implicitly. We propose a representation lear…

Cited by 89PDFScholar
2021

ME-PCN: Point Completion Conditioned on Mask Emptiness

ICCV 2021poster

Point completion refers to completing the missing geometries of an object from incomplete observations. Main-stream methods predict the missing shapes by decoding a global feature learned from the input point cloud, which often leads to deficient results in preserving topology consistency and surfac…

Cited by 27PDFcodeScholar
2021

Preservational Learning Improves Self-Supervised Medical Image Models by Reconstructing Diverse Contexts

ICCV 2021poster

Preserving maximal information is the basic principle of designing self-supervised learning methodologies. To reach this goal, contrastive learning adopts an implicit way which is contrasting image pairs. However, we believe it is not fully optimal to simply use the contrastive estimation for preser…

Cited by 116PDFcodeScholar
2021

Refer-It-in-RGBD: A Bottom-Up Approach for 3D Visual Grounding in RGBD Images

CVPR 2021poster

Grounding referring expressions in RGBD image has been an emerging field. We present a novel task of 3D visual grounding in single-view RGBD image where the referred objects are often only partially scanned due to occlusion. In contrast to previous works that directly generate object proposals for g…

Cited by 43PDFScholar
2021

RfD-Net: Point Scene Understanding by Semantic Instance Reconstruction

CVPR 2021poster

Semantic scene understanding from point clouds is particularly challenging as the points reflect only a sparse set of the underlying 3D geometry. Previous works often convert point cloud into regular grids (e.g. voxels or bird-eye view images), and resort to grid-based convolutions for scene underst…

Cited by 93PDFcodeScholar
2020

Deep Fashion3D: A Dataset and Benchmark for 3D Garment Reconstruction from Single Images

ECCV 2020poster

High-fidelity clothing reconstruction is the key to achieving photorealism in a wide range of applications including human digitization, virtual try-on, etc. Recent advances in learning-based approaches have accomplished unprecedented accuracy in recovering unclothed human shape and pose from single…

2020

FPConv: Learning Local Flattening for Point Convolution

CVPR 2020poster

We introduce FPConv, a novel surface-style convolution operator designed for 3D point cloud analysis. Unlike previous methods, FPConv doesn't require transforming to intermediate representation like 3D grid or graph and directly works on surface geometry of point cloud. To be more specific, for each…

Cited by 187PDFcodeScholar
2020

Peeking into occluded joints: A novel framework for crowd pose estimation

ECCV 2020poster

Although occlusion widely exists in nature and remains a fundamental challenge for pose estimation, existing heatmap-based approaches suffer serious degradation on occlusions. Their intrinsic problem is that they directly localize the joints based on visual information; however, the invisible joints…

2020

Skeleton-bridged Point Completion: From Global Inference to Local Adjustment

NeurIPS 2020poster

Point completion refers to complete the missing geometries of objects from partial point clouds. Existing works usually estimate the missing shape by decoding a latent feature encoded from the input points. However, real-world objects are usually with diverse topologies and surface details, which a…

Cited by 62SourcePDFScholar
2020

Total3DUnderstanding: Joint Layout, Object Pose and Mesh Reconstruction for Indoor Scenes From a Single Image

CVPR 2020oral

Semantic reconstruction of indoor scenes refers to both scene understanding and object reconstruction. Existing works either address one part of this problem or focus on independent objects. In this paper, we bridge the gap between understanding and reconstruction, and propose an end-to-end solution…

Cited by 278PDFcodeScholar
2019

A Skeleton-Bridged Deep Learning Approach for Generating Meshes of Complex Topologies From Single RGB Images

CVPR 2019oral

This paper focuses on the challenging task of learning 3D object surface reconstructions from single RGB images. Existing methods achieve varying degrees of success by using different geometric representations. However, they all have their own drawbacks, and cannot well reconstruct those surfaces of…

Cited by 104PDFScholar
2019

Deep Mesh Reconstruction From Single RGB Images via Topology Modification Networks

ICCV 2019poster

Reconstructing the 3D mesh of a general object from a single image is now possible thanks to the latest advances of deep learning technologies. However, due to the nontrivial difficulty of generating a feasible mesh structure, the state-of-the-art approaches often simplify the problem by learning th…

Cited by 246PDFcodeScholar
2019

Deep Reinforcement Learning of Volume-Guided Progressive View Inpainting for 3D Point Scene Completion From a Single Depth Image

CVPR 2019oral

We present a deep reinforcement learning method of progressive view inpainting for 3D point scene completion under volume guidance, achieving high-quality scene reconstruction from only a single depth image with severe occlusion. Our approach is end-to-end, consisting of three modules: 3D scene volu…

Cited by 55PDFScholar
2019

HEMlets Pose: Learning Part-Centric Heatmap Triplets for Accurate 3D Human Pose Estimation

ICCV 2019poster

Estimating 3D human pose from a single image is a challenging task. This work attempts to address the uncertainty of lifting the detected 2D joints to the 3D space by introducing an intermediate state - Part-Centric Heatmap Triplets (HEMlets), which shortens the gap between the 2D observation and th…

Cited by 164PDFScholar
2017

High-Resolution Shape Completion Using Deep Neural Networks for Global Structure and Local Geometry Inference

ICCV 2017spotlight

We propose a data-driven method for recovering missing parts of 3D shapes. Our method is based on a new deep learning architecture consisting of two sub-networks: a global structure inference network and a local geometry refinement network. The global structure inference network incorporates a long…

Cited by 367PDFScholar