← Search

Chi-Chun Lee

36 accepted papers

2026

Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition

ICASSP 2026oral

Speech Emotion Recognition (SER) is typically trained and evaluated on majority-voted labels, which simplifies benchmarking but masks subjectivity and provides little transparency into why predictions are made. This neglects valid minority annotations and limits interpretability. We propose an expla…

Cited by 0SourcePDFScholar
2025

A Dynamic Edge-Selection Mechanism in HRV Hypergraph Learning for Improved Stress Detection

ICASSP 2025accepted

Studies show that individual attributes such as age and gender significantly influence physiological responses and their correlation with stress, often forming complex and overlapping relationships. These attributes are essential for enhancing physiological signal-based stress detection. Our work le…

Cited by 0SourceScholar
2025

Is It Still Fair? Investigating Gender Fairness in Cross-Corpus Speech Emotion Recognition

ICASSP 2025accepted

Speech emotion recognition (SER) is a vital component in various everyday applications. Cross-corpus SER models are increasingly recognized for their ability to generalize performance. However, concerns arise regarding fairness across demographics in diverse corpora. Existing fairness research often…

Cited by 0SourceScholar
2025

Mask Augmentation For Tumor Classification In Medical Images

ICASSP 2025accepted

Tumor detection and classification in medical images are critical for guiding patient management and treatment decisions. However, accurate segmentation and classification of tumors remain challenging due to their small size relative to the overall image. Existing approaches often face difficulties…

Cited by 0SourceScholar
2025

Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition

ICASSP 2025accepted

Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora, domains, or labels. However, acoustic features are inherently…

Cited by 0SourceScholar
2025

Noise-Robust Speech Emotion Recognition Using Shared Self-Supervised Representations with Integrated Speech Enhancement

ICASSP 2025accepted

Recent studies have demonstrated the effectiveness of fine-tuning self-supervised speech representation models for speech emotion recognition (SER). However, applying SER in real-world environments remains challenging due to pervasive noise. Relying on low-accuracy predictions due to noisy speech ca…

Cited by 0SourceScholar
2025

SocialRecNet: A Multimodal LLM-Based Framework for Assessing Social Reciprocity in Autism Spectrum Disorder

ICASSP 2025accepted

Accurate assessment of social reciprocity is crucial for early diagnosis and intervention in Autism Spectrum Disorder (ASD). Traditional methods, often relying on unimodal data or lacking in cross-modal alignment, do not fully capture the complexity of social reciprocity. To address these limitation…

Cited by 0SourceScholar
2025

Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance

ICASSP 2025accepted

Speech Emotion Recognition (SER) systems rely on speech input and emotional labels annotated by humans. However, various emotion databases collect perceptional evaluations in different ways. For instance, the IEMOCAP dataset uses video clips with sounds for annotators to provide their emotional perc…

Cited by 0SourceScholar
2025

Toward Zero-Shot Speech Emotion Recognition Using LLMs in the Absence of Target Data

ICASSP 2025accepted

In generalized Speech Emotion Recognition (SER), traditional generalization techniques like transfer learning and domain adaptation rely on access to some amount of unlabeled target domain data. However, with increasing privacy concerns, building SER systems under zero-shot scenarios, where no targe…

Cited by 0SourceScholar
2025

Valve Token Masked Autoencoder for Missing Recordings on Cardiac Abnormality Classification

ICASSP 2025accepted

Automated auscultation and cardiovascular screening systems for cardiac abnormalities have received growing interest in clinical applications. Still, they face challenges due to missing or invalid recordings caused by technical issues. To address this, we introduce a novel framework leveraging the m…

Cited by 1SourceScholar
2024

Balancing Speaker-Rater Fairness for Gender-Neutral Speech Emotion Recognition

ICASSP 2024accepted

Speech emotion recognition (SER) adds to the humane aspects of voice technologies to enhance user experiences. The ground truth emotion annotations provided by human raters and attributes related to the speakers themselves arise a compounded fairness issue in SER. While there exist works in fair SER…

Cited by 0SourceScholar
2024

GaP-Aug: Gamma Patch-Wise Correction Augmentation Method for Respiratory Sound Classification

ICASSP 2024accepted

Automated auscultation analysis using electronic stethoscope has received growing interest in clinical applications. Recently, researchers showed successes by using deep learning methods to distinguish between pathological respiratory sound classes. Nevertheless, the challenge persists due to the sc…

Cited by 0SourceScholar
2024

In-The-Wild Physiological-Based Stress Detection Using Federated Strategy

ICASSP 2024accepted

Continuously identifying day-to-day mental stress can be realized by accessing wearable devices to measure physiological indicators. However, the nature of bodily signals raises issues of privacy and data heterogeneity. Recent federated learning scheme provides a promising direction to alleviate the…

Cited by 0SourceScholar
2023

Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion Recognition

ICASSP 2023accepted

Modeling cross-lingual speech emotion recognition (SER) has become more prevalent because of its diverse applications. Existing studies have mostly focused on technical approaches that adapt the feature, domain, or label across languages, without considering in detail the similarities between the la…

Cited by 0SourceScholar
2022

An Audio-Saliency Masking Transformer for Audio Emotion Classification in Movies

ICASSP 2022accepted

The process of perception to affective response of humans is gated by a bottom-up saliency mechanism at the sensory level. In specifics, auditory saliency emphasizes audio segments that need to be attended to cognitively appraise and experience emotion. In this work, inspired by this mechanism, we p…

Cited by 0SourceScholar
2022

Exploiting Annotators' Typed Description of Emotion Perception to Maximize Utilization of Ratings for Speech Emotion Recognition

ICASSP 2022accepted

The decision of ground truth for speech emotion recognition (SER) is still a critical issue in affective computing tasks. Previous studies on emotion recognition often rely on consensus labels after aggregating the classes selected by multiple annotators. It is common for a perceptual evaluation con…

Cited by 0SourceScholar
2020

A Dialogical Emotion Decoder for Speech Motion Recognition in Spoken Dialog

ICASSP 2020accepted

Developing a robust emotion speech recognition (SER) system for human dialog is important in advancing conversational agent design. In this paper, we proposed a novel inference algorithm, a dialogical emotion decoding (DED) algorithm, that treats a dialog as a sequence and consecutively decode the e…

Cited by 26SourceScholar
2020

A Siamese Content-Attentive Graph Convolutional Network for Personality Recognition Using Physiology

ICASSP 2020accepted

Affective multimedia content has long been used as stimulation to study an individual's personality using physiology. In this work, we propose a novel Siamese Content-Attentive Graph Convolutional Network (SCA-GCN) to learn a discriminative physiology representation jointly guided by the actual vide…

Cited by 0SourceScholar
2020

Conditional Domain Adversarial Transfer for Robust Cross-Site ADHD Classification Using Functional MRI

ICASSP 2020accepted

There is a growing number of large scale cross-site database collection of resting-state functional magnetic resonance imaging (rs-fMRI) for studying neurobehavioral diseases, such as ADHD. Although a large amount of data benefits machine learning-based classification methods, the idiosyncratic vari…

Cited by 0SourceScholar
2020

Predicting Performance Outcome with a Conversational Graph Convolutional Network for Small Group Interactions

ICASSP 2020accepted

Studying behaviors of members during small group interaction provides objective insights in improving the efficiency of the decision making process in our daily working life. By introducing the use of the graph structure in modeling the natural inter-member conversational ties during such an interac…

Cited by 0SourceScholar
2019

Adversarially-enriched Acoustic Code Vector Learned from Out-of-context Affective Corpus for Robust Emotion Recognition

ICASSP 2019accepted

Advancement in speech emotion recognition technology has brought tremendous potential in designing human-centered applications across a wide range of scenarios. However, due to the difficulty in obtaining large-scale labeled emotion corpus for every application domains, most of the existing database…

Cited by 0SourceScholar
2019

An Event-contrastive Connectome Network for Automatic Assessment of Individual Face Processing and Memory Ability

ICASSP 2019accepted

Human adapt their behaviors by continuously monitoring one another to function socially in our society. The ability to process face identity from memory is a crucial basic capability. In this work, we propose an event-contrastive connectome network (E-cCN) in representing brain's functional connecti…

Cited by 0SourceScholar
2019

An Interaction-aware Attention Network for Speech Emotion Recognition in Spoken Dialogs

ICASSP 2019accepted

Obtaining robust speech emotion recognition (SER) in scenarios of spoken interactions is critical to the developments of next generation human-machine interface. Previous research has largely focused on performing SER by modeling each utterance of the dialog in isolation without considering the tran…

Cited by 0SourceScholar
2019

Every Rating Matters: Joint Learning of Subjective Labels and Individual Annotators for Speech Emotion Classification

ICASSP 2019accepted

Emotion perception is subjective and vary with respect to each individual due to the natural bias of human, such as gender, culture, and age. Conventionally, emotion recognition relies on the consensus, e.g., majority of annotations (hard label) or the distribution of annotations (soft label), and d…

Cited by 0SourceScholar
2019

Learning Semantic-preserving Space Using User Profile and Multimodal Media Content from Political Social Network

ICASSP 2019accepted

The use of social media in politics has dramatically changed the way campaigns are run and how elected officials interact with their constituents. An advanced algorithm is required to analyze and understand this large amount of heterogeneous social media data to investigate several key issues, such…

Cited by 0SourceScholar
2018

A Triplet-Loss Embedded Deep Regressor Network for Estimating Blood Pressure Changes Using Prosodic Features

ICASSP 2018accepted

Studies have shown that measures of personal physiology, e.g., blood pressure (BP) variation and heart rate variability (HRV), is closely related to a subject's psychological states and are being used regularly to track patients' health conditions in medical settings. The conventional method of moni…

Cited by 0SourceScholar
2018

Integrating Perceivers Neural-Perceptual Responses Using a Deep Voting Fusion Network for Automatic Vocal Emotion Decoding

ICASSP 2018accepted

Understanding neuro-perceptual mechanism of vocal emotion perception continues to be an important research direction not only in advancing scientific knowledge but also in inspiring more robust affective computing technologies. The large variabilities in the manifested fMRI signals among subjects ha…

Cited by 0SourceScholar
2018

Learning Lexical Coherence Representation Using LSTM Forget Gate for Children with Autism Spectrum Disorder During Story-Telling

ICASSP 2018accepted

Inability to carry out cohesive narratives has been identified in children with autism spectrum disorder (ASD). However, deriving cohesion measures is often done using manual labeling or relying on expert-crafted features. In this work, we develop a novel LSTM framework to learn the embedded narrati…

Cited by 0SourceScholar
2017

Fusion of multiple emotion perspectives: Improving affect recognition through integrating cross-lingual emotion information

ICASSP 2017accepted

Developing cross-corpus, cross-domain, and cross-language emotion recognition algorithm has becoming more prevalent recently to ensure the wide applicability of robust emotion recognizer. In this work, we propose a computational framework on fusing multiple emotion perspectives by integrating cross-…

Cited by 0SourceScholar
2016

A Gaussian mixture regression approach toward modeling the affective dynamics between acoustically-derived vocal arousal score (VC-AS) and internal brain fMRI bold signal response

ICASSP 2016accepted

Understanding the underlying neuro-perceptual mechanism of humans' ability to decode emotional content in vocal signal is an important research direction. In this paper, we describe our initial research effort into quantitatively modeling the joint dynamics between measures of vocal arousal and bloo…

Cited by 0SourceScholar
2016

A thin-slice perception of emotion? An information theoretic-based framework to identify locally emotion-rich behavior segments for global affect recognition

ICASSP 2016accepted

Human's judgment has been shown to be thin-sliced in nature, i.e., accurate perception can often be achieved for a short duration of exposure to expressive behaviors. In this work, we develop a mutual information-based framework to select the most emotion-rich 20% of local multimodal behavior segmen…

Cited by 0SourceScholar