← Search

Yuan Zong

14 accepted papers

2025

Enhancing Task-Specific Feature Learning with LLMs for Multimodal Emotion and Intent Joint Understanding

ICASSP 2025accepted

This paper introduces our solution, the Task-Specific Feature Learning (TSFL) method, designed to address the second track of the MEIJU Challenge at ICASSP 2025, namely, Imbalanced Emotion and Intent Recognition (English). The TSFL method incorporates three core components: the use of LLM features t…

Cited by 0SourceScholar
2025

Reliable Learning From LLM Features for Multimodal Emotion and Intent Joint Understanding

ICASSP 2025accepted

This paper describes a Reliable Learning Framework (RLF) for the 1st Multimodal Emotion and Intent Joint Understanding (MEIJU) Challenge at ICASSP 2025. Our proposed RLF includes a Hierarchical Interaction Network and a Reliable Fusion Strategy. The former can excavate emotion and intent cues from t…

Cited by 0SourceScholar
2024

Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition

ICASSP 2024accepted

Cross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This p…

Cited by 0SourceScholar
2024

Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution Adaptation

ICASSP 2024accepted

In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new s…

Cited by 0SourceScholar
2024

PAVITS: Exploring Prosody-Aware VITS for End-to-End Emotional Voice Conversion

ICASSP 2024accepted

In this paper, we propose Prosody-aware VITS (PAVITS) for emotional voice conversion (EVC), aiming to achieve two major objectives of EVC: high content naturalness and high emotional naturalness, which are crucial for meeting the demands of human perception. To improve the content naturalness of con…

Cited by 0SourceScholar
2024

Progressively Learning from Macro-Expressions for Micro-Expression Recognition

ICASSP 2024accepted

Micro-expression (ME) recognition is challenging due to the low-intensity facial motions. An idea to overcome this is learning assisted by macro-expressions (MaEs). However, the intensity gap between MaE and ME is so huge that related works fail to effectively leverage MaE’s assistance in overcoming…

Cited by 0SourceScholar
2024

Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition

ICASSP 2024accepted

Swin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above in…

Cited by 0SourceScholar
2023

CMNet: Contrastive Magnification Network for Micro-Expression Recognition

AAAI 2023technical

Micro-Expression Recognition (MER) is challenging because the Micro-Expressions' (ME) motion is too weak to distinguish. This hurdle can be tackled by enhancing intensity for a more accurate acquisition of movements. However, existing magnification strategies tend to use the features of facial image…

Cited by 5SourcePDFScholar
2023

Deep Implicit Distribution Alignment Networks for cross-Corpus Speech Emotion Recognition

ICASSP 2023accepted

In this paper, we propose a novel deep transfer learning method called deep implicit distribution alignment networks (DIDAN) to deal with cross-corpus speech emotion recognition (SER) problem, in which the labeled training (source) and unlabeled testing (target) speech signals come from different co…

Cited by 0SourceScholar
2023

Learning Attention from Attention: Efficient Self-Refinement Transformer for Face Super-Resolution

IJCAI 2023poster

Recently, Transformer-based architecture has been introduced into face super-resolution task due to its advantage in capturing long-range dependencies. However, these approaches tend to integrate global information in a large searching region, which neglect to focus on the most relevant information…

2022

A Novel Micro-Expression Recognition Approach Using Attention-Based Magnification-Adaptive Networks

ICASSP 2022accepted

Micro-Expression recognition (MER) is a challenging task due to the short duration and low intensity of Micro-Expressions. A popular method to tackle this is magnifying MEs so as to enlarge the expression intensity to make recognition easier. However, the single fixed magnification strategy, widely…

Cited by 0SourceScholar
2021

Cross-Corpus Speech Emotion Recognition Using Joint Distribution Adaptive Regression

ICASSP 2021accepted

In this paper, we focus on the research of cross-corpus speech emotion recognition (SER), in which the training and testing speech signals in cross-corpus SER belong to dierent speech corpus. Due to this fact, mismatched feature distributions may exist between the training and testing speech feature…

Cited by 0SourceScholar
2018

Super Wide Regression Network for Unsupervised Cross-Database Facial Expression Recognition

ICASSP 2018accepted

Unsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature dist…

Cited by 0SourceScholar
2018

Unsupervised Cross-Corpus Speech Emotion Recognition Using Domain-Adaptive Subspace Learning

ICASSP 2018accepted

In this paper, we investigate an interesting problem, i.e., unsupervised cross-corpus speech emotion recognition (SER), in which the training and testing speech signals come from two different speech emotion corpora. Meanwhile, the training speech signals are labeled, while the label information of…

Cited by 0SourceScholar