← Search

Weize Quan

8 accepted papers

2026

FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation

CVPR 2026

Autoregressive models for 3D mesh generation suffer from a fundamental limitation: they flatten meshes into long vertex-coordinate sequences. This results in prohibitive computational costs, hindering the efficient synthesis of high-fidelity geometry. We argue this bottleneck stems from operating at

Cited by 0SourceScholar
2025

Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face Generation

ICASSP 2025accepted

Audio-driven portrait animation is an emerging field in multi-modal generation that aims to create lifelike talking face videos from audio input. While significant progress has been made, accurately modeling the relationship between audio signals and various facial motions, such as head poses and ex…

Cited by 0SourceScholar
2025

GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic Expressions

AAAI 2025technical

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a robust framework GoHD designed to produce highly realistic, ex…

2025

PointCFormer: A Relation-Based Progressive Feature Extraction Network for Point Cloud Completion

AAAI 2025technical

Point cloud completion aims to reconstruct the complete 3D shape from incomplete point clouds, and it is crucial for tasks such as 3D object detection and segmentation. Despite the continuous advances in point cloud analysis techniques, feature extraction methods are still confronted with apparent l…

2024

CMG-Net: Robust Normal Estimation for Point Clouds via Chamfer Normal Distance and Multi-Scale Geometry

AAAI 2024technical

This work presents an accurate and robust method for estimating normals from point clouds. In contrast to predecessor approaches that minimize the deviations between the annotated and the predicted normals directly, leading to direction inconsistency, we first propose a new metric termed Chamfer Nor…

2023

DPE: Disentanglement of Pose and Expression for General Video Portrait Editing

CVPR 2023poster

One-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion and transferred simultaneously. However, the entanglement set…

2022

Bi-Directional Modality Fusion Network For Audio-Visual Event Localization

ICASSP 2022accepted

Audio and visual signals stimulate many audio-visual sensory neurons of persons to generate audio-visual contents, helping humans perceive the world. Most of the existing audio-visual event localization approaches focus on generating audio-visual features by fusing the audio and visual modalities fo…

Cited by 0SourceScholar
2018

Learning 3D Keypoint Descriptors for Non-Rigid Shape Matching

ECCV 2018poster

In this paper, we present a novel deep learning framework that derives discriminative local descriptors for 3D surface shapes. In contrast to previous convolutional neural networks (CNNs) that rely on rendering multi-view images or extracting intrinsic shape properties, we parameterize the multi-sca…

Cited by 49SourcePDFScholar