2022
Is Cross-Attention Preferable to Self-Attention for Multi-Modal Emotion Recognition?
ICASSP 2022accepted
Humans express their emotions via facial expressions, voice intonation and word choices. To infer the nature of the underlying emotion, recognition models may use a single modality, such as vision, audio, and text, or a combination of modalities. Generally, models that fuse complementary information…