Learning Class Prototypes for Visual Emotion Recognition
Jiankun Zhu, Sicheng Zhao, Jing Jiang, Zhaopan Xu, Wenbo Tang, Hongxun Yao
Abstract
Visual emotion recognition (VER), which aims at understanding humans’ emotional reactions toward different visual stimuli, has attracted increasing attention. However, because of the subjectivity and complex nature of emotion, existing VER methods suffer from one or more of the following problems: 1) semantic gap: the large affective gap between visual clues and the emotional expressions; 2) overfitting: the lack of model robustness due to unclear features in the emotional category samples; 3) label ambiguity: the overlap between categories caused by diverse emotional responses. To address these limitations, we present a novel VER method named ProtoEmotion (PoE), exploring discriminative emotional representations by jointly learning prototypes of textual emotional expressions and visual features. Specifically, text prototypes build explicit textual features for each emotion category by extracting prototypes of learnable prompts from multiple aspects, reducing semantic differences. The visual prototypes capture the most defining image features of each category, providing a more robust and discriminative feature representation, while bringing samples closer together to reduce overfitting. In addition, to alleviate the label ambiguity, we propose a label smoothing algorithm based on the prototype distance. Extensive experiments demonstrate the effectiveness of PoE, which outperforms the state-of-the-art by 1.37% on FI and 1.52% on EmotionROI datasets.
BibTeX
@inproceedings{icassp2025_learningclasspro,
title = {Learning Class Prototypes for Visual Emotion Recognition},
author = {Jiankun Zhu and Sicheng Zhao and Jing Jiang and Zhaopan Xu and Wenbo Tang and Hongxun Yao},
booktitle = {ICASSP 2025},
year = {2025}
}