← Search

Carlos Busso

36 accepted papers

2026

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools — From Consensus Learning to Ambiguity-Driven Emotion Reasoning

ICML 2026spotlight

Speech Large Language Models (SLLMs) enable high-level emotion reasoning, but often produce ungrounded, text-biased judgments without verifiable acoustic evidence. In contrast, SSL encoders such as WavLM yield strong acoustic representations yet remain opaque discriminative models that offer limited…

Cited by 0SourceScholar
2026

RankList – a Listwise Preference Learning Framework for Predicting Subjective Preferences

AAAI 2026technical

Preference learning has gained significant attention in tasks involving subjective human judgments, such as speech emotion recognition (SER) and image aesthetic assessment. While pairwise frameworks such as RankNet offer robust modeling of relative preferences, they are inherently limited to local c

Cited by 0SourcePDFScholar
2026

Reasoning Beyond Majority Vote: An Explainable SpeechLM Framework for Speech Emotion Recognition

ICASSP 2026oral

Speech Emotion Recognition (SER) is typically trained and evaluated on majority-voted labels, which simplifies benchmarking but masks subjectivity and provides little transparency into why predictions are made. This neglects valid minority annotations and limits interpretability. We propose an expla…

Cited by 0SourcePDFScholar
2025

Domain-Specific Adaptation in Speech Emotion Recognition Using Emotional Distribution Alignment

ICASSP 2025accepted

This work addresses the challenge of building speech emotion recognition models that generalize effectively across different domains, particularly when only limited target domain data is available with or without emotional label information. Traditional models often struggle with cross-domain perfor…

Cited by 0SourceScholar
2025

Efficient Fusion of Computationally Diverse Modalities Using Chunking and Cross-Attention

ICASSP 2025accepted

Emotion recognition is inherently a multimodal problem. Humans use both audible and visual cues to determine a person’s emotions. There has been extensive improvement in the methods we use to fuse audio and visual representations between two unimodal deep-learning models. However, there is a lack of…

Cited by 0SourceScholar
2025

Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition

ICASSP 2025accepted

Cross-corpus speech emotion recognition (SER) plays a vital role in numerous practical applications. Traditional approaches to cross-corpus emotion transfer often concentrate on adapting acoustic features to align with different corpora, domains, or labels. However, acoustic features are inherently…

Cited by 0SourceScholar
2025

Noise-Robust Speech Emotion Recognition Using Shared Self-Supervised Representations with Integrated Speech Enhancement

ICASSP 2025accepted

Recent studies have demonstrated the effectiveness of fine-tuning self-supervised speech representation models for speech emotion recognition (SER). However, applying SER in real-world environments remains challenging due to pervasive noise. Relying on low-accuracy predictions due to noisy speech ca…

Cited by 0SourceScholar
2024

Generalization of Self-Supervised Learning-Based Representations for Cross-Domain Speech Emotion Recognition

ICASSP 2024accepted

Self-supervised learning (SSL) from unlabelled speech data has revolutionized speech representation learning. Among them, wavLM, wav2vec2, HuBERT, and Data2vec have produced benchmark performances on automatic speech recognition. However, few studies have explored the generalization of SSL-based rep…

Cited by 0SourceScholar
2024

Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition

ICASSP 2024accepted

Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion information is indirectly embedded within speaker embeddings, lea…

Cited by 0SourceScholar
2023

Adapting a Self-Supervised Speech Representation for Noisy Speech Emotion Recognition by Using Contrastive Teacher-Student Learning

ICASSP 2023accepted

Studies have shown high performance in the speech emotion recognition (SER) task by fine-tuning a self-supervised speech representation model. Although this model can provide emotionally discriminative embedding in clean conditions, adapting it to a noisy target environment is still required when de…

Cited by 0SourceScholar
2023

Learning Cross-Modal Audiovisual Representations with Ladder Networks for Emotion Recognition

ICASSP 2023accepted

Representation learning is a challenging, but essential task in audiovisual learning. A key challenge is to generate strong cross-modal representations while still capturing discriminative information contained in unimodal features. Properly capturing this information is important to increase accura…

Cited by 0SourceScholar
2023

Phonetic Anchor-Based Transfer Learning to Facilitate Unsupervised Cross-Lingual Speech Emotion Recognition

ICASSP 2023accepted

Modeling cross-lingual speech emotion recognition (SER) has become more prevalent because of its diverse applications. Existing studies have mostly focused on technical approaches that adapt the feature, domain, or label across languages, without considering in detail the similarities between the la…

Cited by 0SourceScholar
2023

Role of Lexical Boundary Information in Chunk-Level Segmentation for Speech Emotion Recognition

ICASSP 2023accepted

Chunk-level speech emotion recognition (SER) is a common modeling scheme to obtain better recognition performance than sentence-level formulations. A key open question is the role of lexical boundary information in the process of splitting a sentence into small chunks. Is there any benefit in provid…

Cited by 0SourceScholar
2023

Unsupervised Domain Adaptation for Preference Learning Based Speech Emotion Recognition

ICASSP 2023accepted

Retrieving speech samples that have specific expressive content has many applications. It is desirable to build a preference learning framework that ranks speech samples according to emotional attribute values that generalize well to new domains. A popular architecture for preference learning is the…

Cited by 0SourceScholar
2022

Driving Anomaly Detection Using Contrastive Multiview Coding to Interpret Cause of Anomaly

IROS 2022poster

Modern advanced driver assistant systems (ADAS) rely on various types of sensors to monitor the vehicle status, driver's behaviors and road condition. The multimodal systems in the vehicle include sensors, such as accelerometers, pressure sensors, cameras, lidar and radars. When looking at a given s…

Cited by 0SourceScholar
2022

Exploiting Annotators' Typed Description of Emotion Perception to Maximize Utilization of Ratings for Speech Emotion Recognition

ICASSP 2022accepted

The decision of ground truth for speech emotion recognition (SER) is still a critical issue in affective computing tasks. Previous studies on emotion recognition often rely on consensus labels after aggregating the classes selected by multiple annotators. It is common for a perceptual evaluation con…

Cited by 0SourceScholar
2022

Incorporating Gaze Behavior Using Joint Embedding With Scene Context for Driver Takeover Detection

ICASSP 2022accepted

Despite the recent advancement in driver assistance systems, most existing solutions and partial automation systems such as SAE Level 2 driving automation systems assume that the driver is in the loop; the human driver must continuously monitor the driving environment. Frequent transition of maneuve…

Cited by 0SourceScholar
2022

Not All Features are Equal: Selection of Robust Features for Speech Emotion Recognition in Noisy Environments

ICASSP 2022accepted

Speech emotion recognition (SER) system deployed in real-world applications often encounters noisy speech. While most noise compensation techniques consider all acoustic features to have equal impact on the SER model, some acoustic features may be more sensitive to noisy conditions. This paper inves…

Cited by 31SourceScholar
2021

Deepemocluster: a Semi-Supervised Framework for Latent Cluster Representation of Speech Emotions

ICASSP 2021accepted

Semi-supervised learning (SSL) is an appealing approach to resolve generalization problem for speech emotion recognition (SER) systems. By utilizing large amounts of unlabeled data, SSL is able to gain extra information about the prior distribution of the data. Typically, it can lead to better and r…

Cited by 0SourceScholar
2019

Estimation of Gaze Region Using Two Dimensional Probabilistic Maps Constructed Using Convolutional Neural Networks

ICASSP 2019accepted

Predicting the gaze of a user can have important applications in human computer interactions (HCI). They find applications in areas such as social interaction, driver distraction, human robot interaction and education. Appearance based models for gaze estimation have significantly improved due to re…

Cited by 0SourceScholar
2019

Retrieving Speech Samples with Similar Emotional Content Using a Triplet Loss Function

ICASSP 2019accepted

The ability to identify speech with similar emotional content is valuable to many applications, including speech retrieval, surveillance, and emotional speech synthesis. While current formulations in speech emotion recognition based on classification or regression are not appropriate for this task,…

Cited by 0SourceScholar
2018

Novel Realizations of Speech-Driven Head Movements with Generative Adversarial Networks

ICASSP 2018accepted

Head movement is an integral part of face-to-face communications. It is important to investigate methodologies to generate naturalistic movements for conversational agents (CAs). The predominant method for head movement generation is using rules based on the meaning of the message. However, the vari…

Cited by 0SourceScholar
2017

A study of speaker verification performance with expressive speech

ICASSP 2017accepted

Expressive speech introduces variations in the acoustic features affecting the performance of speech technology such as speaker verification systems. It is important to identify the range of emotions for which we can reliably estimate speaker verification tasks. This paper studies the performance of…

Cited by 0SourceScholar
2016

A multimodal analysis of synchrony during dyadic interaction using a metric based on sequential pattern mining

ICASSP 2016accepted

In human-human interaction, people tend to adapt to each other as the conversation progresses, mirroring their intonation, speech rate, fundamental frequency, word selection, hand gestures, and head movements. This phenomenon is known as synchrony, convergence, entrainment, and adaptation. Recent st…

Cited by 0SourceScholar
2016

Automatic composition of broadcast news summaries using rank classifiers trained with acoustic and lexical features

ICASSP 2016accepted

Research on automatic speech summarization typically focuses on optimizing objective evaluation criteria, such as the ROUGE metric, which depend on word and phrase overlaps between automatic and manually generated summary documents. However, the actual quality of the speech summarizer largely depend…

Cited by 0SourceScholar
2016

Tradeoff between quality and quantity of emotional annotations to characterize expressive behaviors

ICASSP 2016accepted

Emotional descriptors collected from perceptual evaluations are important in the study of emotions. Many studies on emotion recognition depend on these labels to train classifiers. The reliability of the emotion descriptors vary with the number and quality of the raters. Conducting perceptual evalua…

Cited by 0SourceScholar