EmoRLTalk: Speech-Driven Emotional Facial Animation With Offline Reinforcement Learning
Gaofeng Liu, Xuetong Li, Ruoyu Gao, Ye Yuan, Jian Liu, Hengsen Li, Hong Huo, Tao Fang
Abstract
In recent years, significant breakthroughs have been made in audio-guided 3D facial animation. However, existing methods mainly focus on lip shape and audio consistency and still face key challenges to achieve alignment between facial emotions and speech emotions. To overcome this limitation, we introduce EmoRLTalk, a novel framework that integrates offline reinforcement learning to implicitly capture the intricate relationship between 3D facial landmarks and blendshape parameters, thereby enhancing the granularity of emotional expression. Furthermore, we harness the strength of conditional diffusion models to synthesize facial motions that are emotionally coherent with the input speech. Additionally, based on the multi-task learning paradigm, we construct a collaborative training framework of a regression main task and a classification sub-task. Specifically, we use emotion classification of blendshape as a sub-task to further improve the model’s ability to express facial emotions. To further enhance system controllability, we integrate the ControlNet module, allowing users to achieve precise facial expression control. Extensive experiments demonstrate that EmoRLTalk achieves superior emotional expressiveness and lip-sync performance compared to previous approaches.
BibTeX
@inproceedings{iros2025_emorltalkspeechd,
title = {EmoRLTalk: Speech-Driven Emotional Facial Animation With Offline Reinforcement Learning},
author = {Gaofeng Liu and Xuetong Li and Ruoyu Gao and Ye Yuan and Jian Liu and Hengsen Li and Hong Huo and Tao Fang},
booktitle = {IROS 2025},
year = {2025}
}