← Search

Zhilei Liu

6 accepted papers

2025

NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis

ICASSP 2025accepted

Talking head synthesis is to synthesize a lip-synchronized talking head video using audio. Recently, the capability of NeRF to enhance the realism and texture details of synthesized talking heads has attracted the attention of researchers. However, most current NeRF methods based on audio are exclus…

Cited by 0SourceScholar
2024

NERF-AD: Neural Radiance Field With Attention-Based Disentanglement For Talking Face Synthesis

ICASSP 2024accepted

Talking face synthesis driven by audio is one of the current research hotspots in the fields of multidimensional signal processing and multimedia. Neural Radiance Field (NeRF) has recently been brought to this research field in order to enhance the realism and 3D effect of the generated faces. Howev…

Cited by 0SourceScholar
2021

Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation Fusion

ICASSP 2021accepted

Due to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy…

Cited by 0SourceScholar
2020

Speech Emotion Recognition with Local-Global Aware Deep Representation Learning

ICASSP 2020accepted

Convolutional neural network (CNN) based deep representation learning methods for speech emotion recognition (SER) have demonstrated great success. The basic design of CNN restricts the ability to model only local information well. Capsule network (CapsNet) can overcome the shortages of CNNs to capt…

Cited by 0SourceScholar
2018

Deep Adaptive Attention for Joint Facial Action Unit Detection and Face Alignment

ECCV 2018poster

Facial action unit (AU) detection and face alignment are two highly correlated tasks since facial landmarks can provide precise AU locations to facilitate the extraction of meaningful local features for AU detection. Most existing AU detection works often treat face alignment as a preprocessing and…

Cited by 223SourcePDFScholar