← Search

Zhuo Su

27 accepted papers

2026

FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty Modeling

CVPR 2026

Full-body motion estimation from sparse VR observations is an inherently under-constrained problem, with only three 6-DoF trackers (HMD and controllers) available to infer a full skeletal pose. To address this ambiguity, we introduce a probabilistic framework that models joint orientations as distri

Cited by 0SourceScholar
2026

FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation

CVPR 2026

We present FlexAvatar, a flexible large reconstruction model for high-fidelity 3D head avatars with detailed dynamic deformation from single or sparse images, without requiring camera poses or expression labels. It leverages a transformer-based reconstruction model with structured head query tokens

Cited by 0SourceScholar
2026

GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction

CVPR 2026

Reconstructing photorealistic and animatable 4D head avatars from a single portrait image remains a fundamental challenge in computer vision. While diffusion models have enabled remarkable progress in image and video generation for avatar reconstruction, existing methods primarily rely on 2D priors

Cited by 0SourceScholar
2026

InterCoser: Interactive 3D Character Creation with Disentangled Fine-Grained Features

AAAI 2026technical

This paper aims to interactively generate and edit disentangled 3D characters based on precise user instructions. Existing methods generate and edit 3D characters via rough and simple editing guidance and entangled representations, making it difficult to achieve precise and comprehensive control ove

Cited by 0SourcePDFScholar
2026

POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling

CVPR 2026

Face relighting aims to synthesize realistic portraits under novel illumination while preserving identity and geometry. However, progress remains constrained by the limited availability of large-scale, physically consistent illumination data. To address this, we introduce POLAR, a large-scale and ph

Cited by 0SourceScholar
2025

CoA: Towards Real Image Dehazing via Compression-and-Adaptation

CVPR 2025poster

Learning-based image dehazing algorithms have shown remarkable success in synthetic domains. However, real image dehazing is still in suspense due to computational resource constraints and the diversity of real-world scenes. Therefore, there is an urgent need for an algorithm that excels in both eff…

2025

EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs

AAAI 2025technical

Egocentric human pose estimation (HPE) using wearable sensors is essential for VR/AR applications. Most methods rely solely on either egocentric-view images or sparse Inertial Measurement Unit (IMU) signals, leading to inaccuracies due to self-occlusion in images or the sparseness and drift of inert…

Cited by 2SourcePDFScholar
2025

EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty Modeling

CVPR 2025poster

Estimating full-body motion using the tracking signals of head and hands from VR devices holds great potential for various applications. However, the sparsity and unique distribution of observations present a significant challenge, resulting in an ill-posed problem with multiple feasible solutions (…

2025

RePerformer: Immersive Human-centric Volumetric Videos from Playback to Photoreal Reperformance

CVPR 2025poster

Human-centric volumetric videos offer immersive free-viewpoint experiences, yet existing methods focus either on replaying general dynamic scenes or animating human avatars, limiting their ability to re-perform general dynamic scenes. In this paper, we present RePerformer, a novel Gaussian-based rep…

Cited by 0SourcePDFScholar
2025

SMGDiff: Soccer Motion Generation using Diffusion Probabilistic Models

ICCV 2025poster

Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate interactions between the player and the ball. In this paper, we introduce SMGDiff, a novel two-stage framework for generat…

Cited by 0SourcePDFScholar
2024

HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations

CVPR 2024poster

It is especially challenging to achieve real-time human motion tracking on a standalone VR Head-Mounted Display (HMD) such as Meta Quest and PICO. In this paper we propose HMD-Poser the first unified approach to recover full-body motions using scalable sparse observations from HMD and body-worn IMUs…

2024

HiFi4G: High-Fidelity Human Performance Rendering via Compact Gaussian Splatting

CVPR 2024poster

We have recently seen tremendous progress in photo-real human modeling and rendering. Yet efficiently rendering realistic human performance and integrating it into the rasterization pipeline remains challenging. In this paper we present HiFi4G an explicit and compact Gaussian-based approach for high…

Cited by 49SourcePDFScholar
2024

HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors

NeurIPS 2024poster

Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly hinder their applications in broader scenarios. To tackle these issues, we present **HumanSplat**, which predicts the 3…

2024

Joint2Human: High-Quality 3D Human Generation via Compact Spherical Embedding of 3D Joints

CVPR 2024poster

3D human generation is increasingly significant in various applications. However the direct use of 2D generative methods in 3D generation often results in losing local details while methods that reconstruct geometry from generated images struggle with global view consistency. In this work we introdu…

Cited by 6SourcePDFScholar
2024

OHTA: One-shot Hand Avatar via Data-driven Implicit Priors

CVPR 2024poster

In this paper we delve into the creation of one-shot hand avatars attaining high-fidelity and drivable hand representations swiftly from a single image. With the burgeoning domains of the digital human the need for quick and personalized hand avatar creation has become increasingly critical. Existin…

Cited by 9SourcePDFScholar
2024

Revitalizing Real Image Deraining via a Generic Paradigm towards Multiple Rainy Patterns

IJCAI 2024poster

Synthetic data-driven methods perform well on image rain removal task, but they still face many challenges in real rainfall scenarios due to the complexity and diversity of rainy patterns. In this paper, we propose a new generic paradigm for real image deraining from the perspective of synthesizing…

Cited by 2SourcePDFScholar
2023

Instant-NVR: Instant Neural Volumetric Rendering for Human-Object Interactions From Monocular RGBD Stream

CVPR 2023poster

Convenient 4D modeling of human-object interactions is essential for numerous applications. However, monocular tracking and rendering of complex interaction scenarios remain challenging. In this paper, we propose Instant-NVR, a neural approach for instant volumetric human-object tracking and renderi…

Cited by 21SourcePDFScholar
2023

Realistic Full-Body Tracking from Sparse Observations via Joint-Level Modeling

ICCV 2023poster

To bridge the physical and virtual worlds for rapidly developed VR/AR applications, the ability to realistically drive 3D full-body avatars is of great significance. Although real-time body tracking with only the head-mounted displays (HMDs) and hand controllers is heavily under-constrained, a caref…

Cited by 28PDFcodeScholar
2022

Dynamic Binary Neural Network by Learning Channel-Wise Thresholds

ICASSP 2022accepted

Binary neural networks (BNNs) constrain weights and activations to +1 or -1 with limited storage and computational cost, which is hardware-friendly for portable devices. Recently, BNNs have achieved remarkable progress and been adopted into various fields. However, the performance of BNNs is sensiti…

Cited by 0SourceScholar
2022

Feature Dense Relevance Network for Single Image Dehazing

IJCAI 2022poster

Existing learning-based dehazing methods do not fully use non-local information, which makes the restoration of seriously degraded region very tough. We propose a novel dehazing network by defining the Feature Dense Relevance module (FDR) and the Shallow Feature Mapping module (SFM). The FDR is defi…

Cited by 4SourcePDFScholar
2022

NeuralHOFusion: Neural Volumetric Rendering Under Human-Object Interactions

CVPR 2022poster

4D modeling of human-object interactions is critical for numerous applications. However, efficient volumetric capture and rendering of complex interaction scenarios, especially from sparse inputs, remain challenging. In this paper, we propose NeuralHOFusion, a neural approach for volumetric human-ob…

Cited by 50PDFScholar
2021

Pixel Difference Networks for Efficient Edge Detection

ICCV 2021poster

Recently, deep Convolutional Neural Networks (CNNs) can achieve human-level performance in edge detection with the rich and abstract edge representation capacities. However, the high performance of CNN based edge detection is achieved with a large pretrained CNN backbone, which is memory and energy…

Cited by 452PDFcodeScholar
2020

Dynamic Group Convolution for Accelerating Convolutional Neural Networks

ECCV 2020poster

Replacing normal convolutions with group convolutions can significantly increase the computational efficiency of modern deep convolutional networks, which has been widely adopted in compact network architecture designs. However, existing group convolutions undermine the original network structures b…

2020

RobustFusion: Human Volumetric Capture with Data-driven Visual Cues using a RGBD Camera

ECCV 2020poster

High-quality and complete 4D reconstruction of human activities is critical for immersive VR/AR experience, but it suffers from inherent self-scanning constraint and consequent fragile tracking under the monocular setting. In this paper, inspired by the huge potential of learning-based human modelin…

Cited by 106SourcePDFScholar
2020

Searching Central Difference Convolutional Networks for Face Anti-Spoofing

CVPR 2020poster

Face anti-spoofing (FAS) plays a vital role in face recognition systems. Most state-of-the-art FAS methods 1) rely on stacked convolutions and expert-designed network, which is weak in describing detailed fine-grained information and easily being ineffective when the environment varies (e.g., differ…

Cited by 620PDFcodeScholar