← Search

Chen Guo

18 accepted papers

2026

FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation

CVPR 2026

We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a transformer-based reconstruction model with structured head query tokens

Cited by 0SourceScholar
2026

RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

CVPR 2026

Reconstructing people, objects, and their interactions in 3D is a long-standing and fundamental goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and objects occlude each other, and camera and object motion entangle

Cited by 0SourcecodeScholar
2025

A Federated Learning-Based Intrusion Detection System for Satellite-Terrestrial Integrated Networks

ICASSP 2025accepted

The emergence of Satellite-Terrestrial Integrated Networks (STIN) has significantly expanded terrestrial network coverage but introduced new security threats. Current Intrusion Detection Systems (IDSs) for STIN mostly consider the distributed nature of satellites, overlooking the computational limit…

Cited by 0SourceScholar
2025

DASSL: Domain Agnostic Self-Supervised Learning with Multiple Missing Information Reconstruction Branches

ICASSP 2025accepted

Self-supervised learning (SSL) is a technique used to learn feature representations from unlabeled data. However, existing SSL frameworks either rely too heavily on domain knowledge due to their design based on feature invariance, leading to a lack of domain transferability, or they are based on aut…

Cited by 0SourceScholar
2025

GPT-NER: Named Entity Recognition via Large Language Models

NAACL 2025findings

Despite the fact that large-scale Language Models (LLM) have achieved SOTA performances on a variety of NLP tasks, its performance on NER is still significantly below supervised baselines. This is due to the gap between the two tasks the NER and LLMs: the former is a sequence labeling task in nature…

2025

PHD: Personalized 3D Human Body Fitting with Point Diffusion

ICCV 2025poster

We introduce PHD, a novel approach for personalized 3D human mesh recovery (HMR) and body fitting that leverages user-specific shape information to improve pose estimation accuracy from videos. Traditional HMR methods are designed to be user-agnostic and optimized for generalization. While these met…

2025

Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning

ACL 2025finding

Packing, initially utilized in the pre-training phase, is an optimization technique designed to maximize hardware resource efficiency by combining different training sequences to fit the model’s maximum input length. Although it has demonstrated effectiveness during pre-training, there remains a lac…

2025

Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior

CVPR 2025poster

We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animation with diverse poses from a monocular video is challenging because the observation of pose diversity and view points is…

Cited by 0SourcePDFScholar
2024

4D-DRESS: A 4D Dataset of Real-World Human Clothing With Semantic Annotations

CVPR 2024highlight

The studies of human clothing for digital avatars have predominantly relied on synthetic datasets. While easy to collect synthetic data often fall short in realism and fail to capture authentic clothing dynamics. Addressing this gap we introduce 4D-DRESS the first real-world 4D dataset advancing hum…

2024

HSR: Holistic 3D Human-Scene Reconstruction from Monocular Videos

ECCV 2024poster

"An overarching goal for computer-aided perception systems is the holistic understanding of the human-centric 3D world, including faithful reconstructions of humans, scenes, and their global spatial relationships. While recent progress in monocular 3D reconstruction has been made for footage of eith…

Cited by 3SourcePDFScholar
2024

MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild

CVPR 2024poster

We present MultiPly a novel framework to reconstruct multiple people in 3D from monocular in-the-wild videos. Reconstructing multiple individuals moving and interacting naturally from monocular in-the-wild videos poses a challenging task. Addressing it necessitates precise pixel-level disentanglemen…

Cited by 9SourcePDFScholar
2023

EMDB: The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild

ICCV 2023poster

We present EMDB, the Electromagnetic Database of Global 3D Human Pose and Shape in the Wild. EMDB is a novel dataset that contains high-quality 3D SMPL pose and shape parameters with global body and camera trajectories for in-the-wild videos. We use body-worn, wireless electromagnetic (EM) sensors a…

Cited by 52PDFcodeScholar
2023

Gender-Cartoon: Image Cartoonization Method Based on Gender Classification

ICASSP 2023accepted

Qin Opera art is one of China’s intangible cultural heritage, and its influence is gradually declining. The cartoonization of Qin Opera is one of the feasible methods. However, current cartoonization methods suffer from the inability to classify and accurately cartoonize Qinqiang portraits by gender…

Cited by 0SourceScholar
2023

Hi4D: 4D Instance Segmentation of Close Human Interaction

CVPR 2023poster

We propose Hi4D, a method and dataset for the auto analysis of physically close human-human interaction under prolonged contact. Robustly disentangling several in-contact subjects is a challenging task due to occlusions and complex shapes. Hence, existing multi-view systems typically fuse 3D surface…

2023

Vid2Avatar: 3D Avatar Reconstruction From Videos in the Wild via Self-Supervised Scene Decomposition

CVPR 2023poster

We present Vid2Avatar, a method to learn human avatars from monocular in-the-wild videos. Reconstructing humans that move naturally from monocular in-the-wild videos is difficult. Solving it requires accurately separating humans from arbitrary backgrounds. Moreover, it requires reconstructing detail…

2023

X-Avatar: Expressive Human Avatars

CVPR 2023poster

We present X-Avatar, a novel avatar model that captures the full expressiveness of digital humans to bring about life-like experiences in telepresence, AR/VR and beyond. Our method models bodies, hands, facial expressions and appearance in a holistic fashion and can be learned from either full 3D sc…

2022

PINA: Learning a Personalized Implicit Neural Avatar From a Single RGB-D Video Sequence

CVPR 2022poster

We present a novel method to learn Personalized Implicit Neural Avatars (PINA) from a short RGB-D sequence. This allows non-expert users to create a detailed and personalized virtual copy of themselves, which can be animated with realistic clothing deformations. PINA does not require complete scans,…

Cited by 73PDFScholar