← Search

Dong-Ming Yan

15 accepted papers

2026

FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation

CVPR 2026

Autoregressive models for 3D mesh generation suffer from a fundamental limitation: they flatten meshes into long vertex-coordinate sequences. This results in prohibitive computational costs, hindering the efficient synthesis of high-fidelity geometry. We argue this bottleneck stems from operating at

Cited by 0SourceScholar
2025

Diffused Poses and Distilled Expressions for Controllable Audio-driven Talking Face Generation

ICASSP 2025accepted

Audio-driven portrait animation is an emerging field in multi-modal generation that aims to create lifelike talking face videos from audio input. While significant progress has been made, accurately modeling the relationship between audio signals and various facial motions, such as head poses and ex…

Cited by 0SourceScholar
2025

GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic Expressions

AAAI 2025technical

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a robust framework GoHD designed to produce highly realistic, ex…

2025

Occlusion-aware Non-Rigid Point Cloud Registration via Unsupervised Neural Deformation Correntropy

ICLR 2025poster

Non-rigid alignment of point clouds is crucial for scene understanding, reconstruction, and various computer vision and robotics tasks. Recent advancements in implicit deformation networks for non-rigid registration have significantly reduced the reliance on large amounts of annotated training data.…

2025

PointCFormer: A Relation-Based Progressive Feature Extraction Network for Point Cloud Completion

AAAI 2025technical

Point cloud completion aims to reconstruct the complete 3D shape from incomplete point clouds, and it is crucial for tasks such as 3D object detection and segmentation. Despite the continuous advances in point cloud analysis techniques, feature extraction methods are still confronted with apparent l…

2025

Revisiting CAD Model Generation by Learning Raster Sketch

AAAI 2025technical

The integration of deep generative networks into generating Computer-Aided Design (CAD) models has garnered increasing attention over recent years. Traditional methods often rely on discrete sequences of parametric line/curve segments to represent sketches. Differently, we introduce RECAD, a novel f…

Cited by 0SourcePDFScholar
2024

CMG-Net: Robust Normal Estimation for Point Clouds via Chamfer Normal Distance and Multi-Scale Geometry

AAAI 2024technical

This work presents an accurate and robust method for estimating normals from point clouds. In contrast to predecessor approaches that minimize the deviations between the annotated and the predicted normals directly, leading to direction inconsistency, we first propose a new metric termed Chamfer Nor…

2024

Correspondence-Free Non-Rigid Point Set Registration Using Unsupervised Clustering Analysis

CVPR 2024highlight

This paper presents a novel non-rigid point set registration method that is inspired by unsupervised clustering analysis. Unlike previous approaches that treat the source and target point sets as separate entities we develop a holistic framework where they are formulated as clustering centroids and…

2023

DPE: Disentanglement of Pose and Expression for General Video Portrait Editing

CVPR 2023poster

One-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion and transferred simultaneously. However, the entanglement set…

2023

Structure-Aware Surface Reconstruction via Primitive Assembly

ICCV 2023poster

We propose a novel and efficient method for reconstructing manifold surfaces from point clouds. Unlike previous approaches that use dense implicit reconstructions or piecewise approximations and overlook inherent structures like quadrics in CAD models, our method faithfully preserves these quadric s…

Cited by 4PDFScholar
2022

Bi-Directional Modality Fusion Network For Audio-Visual Event Localization

ICASSP 2022accepted

Audio and visual signals stimulate many audio-visual sensory neurons of persons to generate audio-visual contents, helping humans perceive the world. Most of the existing audio-visual event localization approaches focus on generating audio-visual features by fusing the audio and visual modalities fo…

Cited by 0SourceScholar
2022

GraphFit: Learning Multi-Scale Graph-Convolutional Representation for Point Cloud Normal Estimation

ECCV 2022poster

"We propose a precise and efficient normal estimation method that can deal with noise and nonuniform density for unstructured 3D point clouds. Unlike existing approaches that directly take patches and ignore the local neighborhood relationships, which make them susceptible to challenging regions suc…

2021

LARNet: Lie Algebra Residual Network for Face Recognition

ICML 2021spotlight

Face recognition is an important yet challenging problem in computer vision. A major challenge in practical face recognition applications lies in significant variations between profile and frontal faces. Traditional techniques address this challenge either by synthesizing frontal faces or by pose in…

Cited by 33SourcePDFScholar
2019

A Robust Local Spectral Descriptor for Matching Non-Rigid Shapes With Incompatible Shape Structures

CVPR 2019poster

Constructing a robust and discriminative local descriptor for 3D shape is a key component of many computer vision applications. Although existing learning-based approaches can achieve good performance in some specific benchmarks, they usually fail to learn enough information from shapes with differe…

Cited by 25PDFScholar
2018

Learning 3D Keypoint Descriptors for Non-Rigid Shape Matching

ECCV 2018poster

In this paper, we present a novel deep learning framework that derives discriminative local descriptors for 3D surface shapes. In contrast to previous convolutional neural networks (CNNs) that rely on rendering multi-view images or extracting intrinsic shape properties, we parameterize the multi-sca…

Cited by 49SourcePDFScholar