2025
Efficient Fusion of Computationally Diverse Modalities Using Chunking and Cross-Attention
ICASSP 2025accepted
Emotion recognition is inherently a multimodal problem. Humans use both audible and visual cues to determine a person’s emotions. There has been extensive improvement in the methods we use to fuse audio and visual representations between two unimodal deep-learning models. However, there is a lack of…