← Search

Wenxiong Kang

12 accepted papers

2026

Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network

AAAI 2026technical

Garment simulation is fundamental to various applications in computer vision and graphics, from virtual try-on to digital human modelling. However, conventional physics-based methods remain computationally expensive, hindering their application in time-sensitive scenarios. While graph neural network

Cited by 0SourcePDFScholar
2025

Diffusion-Guided Graph Data Augmentation

NeurIPS 2025poster

Graph Neural Networks (GNNs) have achieved remarkable success in a wide range of applications. However, when trained on limited or low-diversity datasets, GNNs are prone to overfitting and memorization, which impacts their generalization. To address this, graph data augmentation (GDA) has become a c…

Cited by 0SourceScholar
2025

GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer

IJCAI 2025

Speech-driven talking head generation is a critical yet challenging task with applications in augmented reality and virtual human modeling. While recent approaches using autoregressive and diffusion-based models have achieved notable progress, they often suffer from modality inconsistencies, particu

Cited by 0SourcePDFScholar
2024

FAVOR: Full-Body AR-Driven Virtual Object Rearrangement Guided by Instruction Text

AAAI 2024technical

Rearrangement operations form the crux of interactions between humans and their environment. The ability to generate natural, fluid sequences of this operation is of essential value in AR/VR and CG. Bridging a gap in the field, our study introduces FAVOR: a novel dataset for Full-body AR-driven Virt…

2024

Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial Animation

AAAI 2024technical

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while neglecting to model the subject-specific speaking style, of…

2023

CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo

IJCAI 2023poster

The core of Multi-view Stereo(MVS) is the matching process among reference and source pixels. Cost aggregation plays a significant role in this process, while previous methods focus on handling it via CNNs. This may inherit the natural limitation of CNNs that fail to discriminate repetitive or incor…

Cited by 18SourcePDFScholar
2023

PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-Modal Distillation and Super-Voxel Clustering

ICCV 2023poster

Semantic segmentation of point clouds usually requires exhausting efforts of human annotations, hence it attracts wide attention to a challenging topic of learning from unlabeled or weaker form of annotations. In this paper, we take the first attempt for fully unsupervised semantic segmentation of p…

Cited by 10PDFcodeScholar
2022

On the Importance of Different Frequency Bins for Speaker Verification

ICASSP 2022accepted

The majority of modern speaker verification systems take spectral analysis-based features as input, which contains multiple frequency bins. Naturally, there would be a question of whether all different frequency bins contribute equally to the speaker verification system performance? In this paper, w…

Cited by 0SourceScholar
2021

Digging Into Uncertainty in Self-Supervised Multi-View Stereo

ICCV 2021poster

Self-supervised Multi-view stereo (MVS) with a pretext task of image reconstruction has achieved significant progress recently. However, previous methods are built upon intuitions, lacking comprehensive explanations about the effectiveness of the pretext task in self-supervised MVS. To this end, we…

Cited by 70PDFcodeScholar
2021

Self-supervised Multi-view Stereo via Effective Co-Segmentation and Data-Augmentation

AAAI 2021technical

Recent studies have witnessed that self-supervised methods based on view synthesis obtain clear progress on multi-view stereo (MVS). However, existing methods rely on the assumption that the corresponding points among different views share the same color, which may not always be true in practice. Th…

2020

Dynamic Group Convolution for Accelerating Convolutional Neural Networks

ECCV 2020poster

Replacing normal convolutions with group convolutions can significantly increase the computational efficiency of modern deep convolutional networks, which has been widely adopted in compact network architecture designs. However, existing group convolutions undermine the original network structures b…

2020

JGR-P2O: Joint Graph Reasoning based Pixel-to-Offset Prediction Network for 3D Hand Pose Estimation from a Single Depth Image

ECCV 2020poster

State-of-the-art single depth image-based 3D hand pose estimation methods are based on dense predictions, including voxel-to-voxel predictions, point-to-point regression, and pixel-wise estimations. Despite the good performance, those methods have a few issues in nature, such as the poor trade-off b…