Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction Generation
Qincheng Lv, Xiaofeng Liu, Jie Li, Rongrong Ni, Pujun Xue, Siyang Song
Abstract
Facial reactions convey crucial emotional information and coordinating interpersonal relationships in human dyadic interactions. While existing Multiple Appropriate Facial Reaction Generation (MAFRG) methods focus on generating multiple reasonable facial reactions, none of these approaches combines 2D and 3D facial behaviour information nor account for the influence of individuals’ facial identities, leading to inconsistencies in the generated facial reactions and limited capability in capturing subtle variations in facial depth and expression dynamics. This paper proposes a novel Hierarchical Multimodal Decoupling-Fusion (HMDF) framework that decouples 3D facial identity from expression behaviors, eliminating identity-based interference in the reaction generation process, which are integrated with audio-visual features through a cross-attention mechanism. Experiments show that our framework achieved the enhanced diversity and synchrony in the generated facial reactions.
BibTeX
@inproceedings{icassp2025_hierarchicalmult,
title = {Hierarchical Multimodal Decoupling-Fusion Framework for offline Multiple Appropriate Facial Reaction Generation},
author = {Qincheng Lv and Xiaofeng Liu and Jie Li and Rongrong Ni and Pujun Xue and Siyang Song},
booktitle = {ICASSP 2025},
year = {2025}
}