← Search

Linchao Bao

21 accepted papers

2025

DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh

CVPR 2025poster

Text-driven avatar generation has gained significant attention owing to its convenience. However, existing methods typically model the human body with all garments as a single 3D model, limiting its usability, such as clothing replacement, and reducing user control over the generation process. To ov…

Cited by 2SourcePDFScholar
2025

MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head Editing

CVPR 2025poster

Creating high-fidelity head avatars from multi-view videos is essential for many AR/VR applications. However, current methods often struggle to achieve high-quality renderings across all head components (e.g., skin vs. hair) due to the limitations of using one single representation for elements with…

2024

RePOSE: 3D Human Pose Estimation via Spatio-Temporal Depth Relational Consistency

ECCV 2024poster

"We introduce RePOSE, a simple yet effective approach for addressing occlusion challenges in the learning of 3D human pose estimation (HPE) from videos. Conventional approaches typically employ absolute depth signals as supervision, which are adept at discernible keypoints but become less reliable w…

2023

FFHQ-UV: Normalized Facial UV-Texture Dataset for 3D Face Reconstruction

CVPR 2023poster

We present a large-scale facial UV-texture dataset that contains over 50,000 high-quality texture UV-maps with even illuminations, neutral expressions, and cleaned facial regions, which are desired characteristics for rendering realistic 3D face models under different lighting conditions. The datase…

2023

Get3DHuman: Lifting StyleGAN-Human into a 3D Generative Model Using Pixel-Aligned Reconstruction Priors

ICCV 2023poster

Fast generation of high-quality 3D digital humans is important to a vast number of applications ranging from entertainment to professional concerns. Recent advances in differentiable rendering have enabled the training of 3D generative models without requiring 3D ground truths. However, the quality…

Cited by 24PDFScholar
2023

Skinned Motion Retargeting With Residual Perception of Motion Semantics & Geometry

CVPR 2023poster

A good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which relies on two neural modification modules, to adjust the sou…

2022

CAR: Class-Aware Regularizations for Semantic Segmentation

ECCV 2022poster

"Recent segmentation methods, such as OCR and CPNet, utilizing “class level” information in addition to pixel features, have achieved notable success for boosting the accuracy of existing network modules. However, the extracted class-level information was simply concatenated to pixel features, witho…

2022

REALY: Rethinking the Evaluation of 3D Face Reconstruction

ECCV 2022poster

"The evaluation of 3D face reconstruction results typically relies on a rigid shape alignment between the estimated 3D model and the ground-truth scan. We observe that aligning two shapes with different reference points can largely affect the evaluation results. This poses difficulties for precisely…

2021

Audio2Gestures: Generating Diverse Gestures From Speech Audio With Conditional Variational Autoencoders

ICCV 2021poster

Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping between audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all possible target motions, resulting in plain/boring motions during…

Cited by 130PDFcodeScholar
2021

Model-Based 3D Hand Reconstruction via Self-Supervised Learning

CVPR 2021poster

Reconstructing a 3D hand from a single-view RGB image is challenging due to various hand configurations and depth ambiguity. To reliably reconstruct a 3D hand from a monocular image, most state-of-the-art methods heavily rely on 3D annotations at the training stage, but obtaining 3D annotations is e…

Cited by 124PDFcodeScholar
2021

Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation

CVPR 2021poster

Image-to-Image (I2I) multi-domain translation models are usually evaluated also using the quality of their semantic interpolation results. However, state-of-the-art models frequently show abrupt changes in the image appearance during interpolation, and usually perform poorly in interpolations across…

Cited by 58PDFScholar
2019

Face Anti-Spoofing: Model Matters, so Does Data

CVPR 2019poster

Face anti-spoofing is an important task in full-stack face applications including face detection, verification, and recognition. Previous approaches build models on datasets which do not simulate the real-world data well (e.g., small scale, insignificant variance, etc.). Existing models may rely on…

Cited by 302PDFScholar
2019

MHP-VOS: Multiple Hypotheses Propagation for Video Object Segmentation

CVPR 2019oral

We address the problem of semi-supervised video object segmentation (VOS), where the masks of objects of interests are given in the first frame of an input video. To deal with challenging cases where objects are occluded or missing, previous work relies on greedy data association strategies that mak…

Cited by 67PDFcodeScholar
2019

MVF-Net: Multi-View 3D Face Morphable Model Regression

CVPR 2019poster

We address the problem of recovering the 3D geometry of a human face from a set of facial images in multiple views. While recent studies have shown impressive progress in 3D Morphable Model (3DMM) based facial reconstruction, the settings are mostly restricted to a single view. There is an inherent…

Cited by 144PDFScholar
2019

Self-Supervised Spatio-Temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics

CVPR 2019poster

We address the problem of video representation learning without human-annotated labels. While previous efforts address the problem by designing novel self-supervised tasks using video data, the learned features are merely on a frame-by-frame basis, which are not applicable to many video analytic tas…

Cited by 259PDFcodeScholar
2018

CNN in MRF: Video Object Segmentation via Inference in a CNN-Based Higher-Order Spatio-Temporal MRF

CVPR 2018poster

This paper addresses the problem of video object segmentation, where the initial object mask is given in the first frame of an input video. We propose a novel spatio-temporal Markov Random Field (MRF) model defined over pixels to handle this problem. Unlike conventional MRF models, the spatial depen…

Cited by 209SourcePDFScholar
2018

Dynamic Scene Deblurring Using Spatially Variant Recurrent Neural Networks

CVPR 2018poster

Due to the spatially variant blur caused by camera shake and object motions under different scene depths, deblurring images captured from dynamic scenes is challenging. Although recent works based on deep neural networks have shown great progress on this problem, their models are usually large and c…

Cited by 466SourcePDFScholar
2018

Modeling Varying Camera-IMU Time Offset in Optimization-Based Visual-Inertial Odometry

ECCV 2018poster

Combining cameras and inertial measurement units (IMUs) has been proven effective in motion tracking, as these two sensing modalities offer complementary characteristics that are suitable for fusion. While most works focus on global-shutter cameras and synchronized sensor measurements, consumer-grad…

Cited by 25SourcePDFScholar
2018

VITAL: VIsual Tracking via Adversarial Learning

CVPR 2018poster

The tracking-by-detection framework consists of two stages, i.e., drawing samples around the target object in the first stage and classifying each sample as the target object or as background in the second stage. The performance of existing tracking-by-detection trackers using deep classification ne…

Cited by 654SourcePDFScholar