← Search

Björn W. Schuller

61 accepted papers

2025

DFingerNet: Noise-Adaptive Speech Enhancement for Hearing Aids

ICASSP 2025accepted

The DeepFilterNet (DFN) architecture was recently proposed as a deep learning model suited for hearing aid devices. Despite its competitive performance on numerous benchmarks, it still follows a ‘one-size-fits-all’ approach, which aims to train a single, monolithic architecture that generalises acro…

Cited by 0SourceScholar
2025

Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models

ICASSP 2025accepted

While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still remains a significant challenge. In this paper, we introduce ParaEVITS, a novel emotional TTS framework that leverages the…

Cited by 0SourceScholar
2025

GNCL: A Graph Neural Network with Consistency Loss for Segment-Level Spoofed Speech Detection

ICASSP 2025accepted

Segment-level spoofed speech detection focuses on recognizing fake or synthetic segments within identifying partially spoofed speech. Nevertheless, existing models for this segment-level task usually overlook latent local relationships between fake and bona fide segments, and further, a lack of inte…

Cited by 0SourceScholar
2025

MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding Challenge

ICASSP 2025accepted

Multimodal Emotion and Intent Joint Understanding (MEIJU) aims to decode the semantic information expressed in the multimodal dialogues while inferring the emotions and intents, providing users with a more humanized human-machine interaction experience. However, challenges such as difficulties in da…

Cited by 0SourceScholar
2025

SSE: A Speaking Style Extractor Based on Fine-Grained Contrastive Learning between Speech and Descriptive Text

ICASSP 2025accepted

Effective extraction of paralinguistic features from speech, such as emotion, accent, and age, remains a challenging task in speech processing. Traditional methods typically address each type of paralinguistic information with separate classification or regression tasks—e. g., emotion recognition, a…

Cited by 0SourceScholar
2025

XDGesture: An xLSTM-based Diffusion Model for Co-speech Gesture Generation

ICASSP 2025accepted

In multimodal human-computer interaction, generating co-speech gestures is crucial for enhancing interaction naturalness and user experience. However, achieving synchronized and natural gesture sequences remains a significant challenge due to the complexity of modeling temporal dependencies across d…

Cited by 0SourceScholar
2024

Customising General Large Language Models for Specialised Emotion Recognition Tasks

ICASSP 2024accepted

The advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning,…

Cited by 0SourceScholar
2024

Deep Fusion of Shifted MLP and CNN for Medical Image Segmentation

ICASSP 2024accepted

Medical image segmentation is an important task in modern analysis of medical images. Current methods tend to extract either local features with convolutions or global features with Transformers. However, few of them are able to effectively fuse global and local features to facilitate segmentation.…

Cited by 0SourceScholar
2024

Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition

ICASSP 2024accepted

Cross-corpus speech emotion recognition (SER) aims to transfer emotional knowledge from a labeled source corpus to an unlabeled corpus. However, prior methods require access to source data during adaptation, which is unattainable in real-life scenarios due to data privacy protection concerns. This p…

Cited by 0SourceScholar
2024

Exploring Meta Information for Audio-Based Zero-Shot Bird Classification

ICASSP 2024accepted

Advances in passive acoustic monitoring and machine learning have led to the procurement of vast datasets for computational bioacoustic research. Nevertheless, data scarcity is still an issue for rare and underrepresented species. This study investigates how meta-information can improve zero-shot au…

Cited by 8SourceScholar
2024

HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech

ICASSP 2024accepted

Automatically detecting Alzheimer’s Disease (AD) from spontaneous speech plays an important role in its early diagnosis. Recent approaches highly rely on the Transformer architectures due to its efficiency in modelling long-range context dependencies. However, the quadratic increase in computational…

Cited by 0SourceScholar
2024

Improving Speaker-Independent Speech Emotion Recognition using Dynamic Joint Distribution Adaptation

ICASSP 2024accepted

In speaker-independent speech emotion recognition, the training and testing samples are collected from diverse speakers, leading to a multi-domain shift challenge across the feature distributions of data from different speakers. Consequently, when the trained model is confronted with data from new s…

Cited by 0SourceScholar
2024

Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty Estimation

ICASSP 2024accepted

Heart murmurs are a common manifestation of cardiovascular diseases and can provide crucial clues to early cardiac abnormalities. While most current research methods primarily focus on the accuracy of models, they often overlook other important aspects such as the interpretability of machine learnin…

Cited by 0SourceScholar
2024

Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition

ICASSP 2024accepted

Swin-Transformer has demonstrated remarkable success in computer vision by leveraging its hierarchical feature representation based on Transformer. In speech signals, emotional information is distributed across different scales of speech features, e. g., word, phrase, and utterance. Drawing above in…

Cited by 0SourceScholar
2024

Synthia's Melody: A Benchmark Framework for Unsupervised Domain Adaptation in Audio

ICASSP 2024accepted

Despite significant advancements in deep learning for vision and natural language, unsupervised domain adaptation in audio remains relatively unexplored. We, in part, attribute this to the lack of an appropriate benchmark dataset. To address this gap, we present Synthia’s melody, a novel audio data…

Cited by 0SourceScholar
2024

Task Selection and Assignment for Multi-Modal Multi-Task Dialogue Act Classification with Non-Stationary Multi-Armed Bandits

ICASSP 2024accepted

Multi-task learning (MTL) aims to improve the performance of a primary task by jointly learning with related auxiliary tasks. Traditional MTL methods select tasks randomly during training. However, both previous studies and our results suggest that such a random selection of tasks may not be helpful…

Cited by 1SourceScholar
2023

Audio Barlow Twins: Self-Supervised Audio Representation Learning

ICASSP 2023accepted

The Barlow Twins self-supervised learning objective requires neither negative samples or asymmetric learning updates, achieving results on a par with the current state-of-the-art within Computer Vision. As such, we present Audio Barlow Twins, a novel self-supervised audio representation learning app…

Cited by 0SourceScholar
2023

Daily Mental Health Monitoring from Speech: A Real-World Japanese Dataset and Multitask Learning Analysis

ICASSP 2023accepted

Translating mental health recognition from clinical research into real-world application requires extensive data, yet existing emotion datasets are impoverished in terms of daily mental health monitoring, especially when aiming for self-reported anxiety and depression recognition. We introduce the J…

Cited by 0SourceScholar
2023

Fast Yet Effective Speech Emotion Recognition with Self-Distillation

ICASSP 2023accepted

Speech emotion recognition (SER) is the task of recognising humans’ emotional states from speech. SER is extremely prevalent in helping dialogue systems to truly understand our emotions and become a trustworthy human conversational partner. Due to the lengthy nature of speech, SER also suffers from…

Cited by 0SourceScholar
2023

Federated Intelligent Terminals Facilitate Stuttering Monitoring

ICASSP 2023accepted

Stuttering is a complicated language disorder. The most common form of stuttering is developmental stuttering, which begins in childhood. Early monitoring and intervention are essential for the treatment of children with stuttering. Automatic speech recognition technology has shown its great potenti…

Cited by 0SourceScholar
2023

Hearttoheart: The Arts of Infant Versus Adult-Directed Speech Classification

ICASSP 2023accepted

Psycholinguistics researchers investigate child language exposure by studying children’s language environment. A main factor is whether, in humanistic heart-to-heart dialogue, the speech is directed to the infant (infant-directed speech) versus to another adult (adult-directed speech). The former ha…

Cited by 3SourceScholar
2023

Hierarchical Network with Decoupled Knowledge Distillation for Speech Emotion Recognition

ICASSP 2023accepted

The goal of Speech Emotion Recognition (SER) is to enable computers to recognize the emotion category of a given utterance in the same way that humans do. The accuracy of SER is strongly dependent on the validity of the utterance-level representation obtained by the model. Nevertheless, the "dark kn…

Cited by 0SourceScholar
2023

Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured Learning

ICASSP 2023accepted

Speech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep learning has been investigated to improve the performance o…

Cited by 0SourceScholar
2023

Large-Scale Nonverbal Vocalization Detection Using Transformers

ICASSP 2023accepted

Detecting emotionally expressive nonverbal vocalizations is essential to developing technologies that can converse fluently with humans. The affective computing community has largely focused on understanding the intonation of emotional speech and language. However, advances in the study of vocal emo…

Cited by 18SourceScholar
2023

Masking Speech Contents by Random Splicing: is Emotional Expression Preserved?

ICASSP 2023accepted

We discuss the influence of random splicing on the perception of emotional expression in speech signals. Random splicing is the randomized reconstruction of short audio snippets with the aim to obfuscate the speech contents. A part of the German parliament recordings has been random spliced and both…

Cited by 0SourceScholar
2023

Positive-Pair Redundancy Reduction Regularisation for Speech-Based Asthma Diagnosis Prediction

ICASSP 2023accepted

Asthma affects an estimated 334 million people worldwide, causing over 461 000 deaths. Exacerbations or asthma attacks can be predicted with new sensor technologies. We explore how recordings of human voice, and machine learning can provide better diagnostics for pulmonary diseases like asthma, as w…

Cited by 6SourceScholar
2023

Zero-Shot Speech Emotion Recognition Using Generative Learning with Reconstructed Prototypes

ICASSP 2023accepted

Zero-shot Speech Emotion Recognition (SER) enables machines to perceive unseen-emotional speech without knowing any samples from these emotional states, which is helpful in audio-based autonomous affective computing. However, existing works on zero-shot SER directly employ original prototypes and on…

Cited by 7SourceScholar
2022

A Glance-and-Gaze Network for Respiratory Sound Classification

ICASSP 2022accepted

A plethora of great successes has been achieved by the existing convolutional neural networks (CNN) for respiratory sound classification. Nevertheless, simultaneously capturing both the local and global features can never be an easy task due to the limitation of a CNN’s structure. In this contributi…

Cited by 8SourceScholar
2022

An Overview of the FIRST ICASSP Special Session on Computer Audition for Healthcare

ICASSP 2022accepted

Audio has been increasingly used as a novel digital phenotype that carries important information of the subject’s health status. We can find tremendous efforts given to this young and promising field, i.e., computer audition for healthcare (CA4H), whereas the application scenarios have not been full…

Cited by 0SourceScholar
2022

Convoluational Transformer With Adaptive Position Embedding For Covid-19 Detection From Cough Sounds

ICASSP 2022accepted

Covid-19 has caused a huge health crisis worldwide in the past two years. Although an early detection of the virus through nucleic acid screening can considerably reduce its spread, the efficiency of this diagnostic process is limited by its complexity and costs. Hence, an effective and inexpensive…

Cited by 7SourceScholar
2021

A Novel Attention-Based Gated Recurrent Unit and its Efficacy in Speech Emotion Recognition

ICASSP 2021accepted

Notwithstanding the significant advancements in the field of deep learning, the basic long short-term memory (LSTM) or Gated Recurrent Unit (GRU) units have largely remained unchanged and unexplored. There are several possibilities in advancing the state-of-art by rightly adapting and enhancing the…

Cited by 0SourceScholar
2021

Hierarchical Attention-Based Temporal Convolutional Networks for Eeg-Based Emotion Recognition

ICASSP 2021accepted

EEG-based emotion recognition is an effective way to infer the inner emotional state of human beings. Recently, deep learning methods, particularly long short-term memory recurrent neural networks (LSTM-RNNs), have made encouraging progress for in the field of emotion recognition. However, the LSTM-…

Cited by 34SourceScholar
2021

Speech Emotion Recognition Using Semantic Information

ICASSP 2021accepted

Speech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of Deep Neural Networks (DNN), most of the studies in the literatu…

Cited by 0SourceScholar
2021

The Role of Task and Acoustic Similarity in Audio Transfer Learning: Insights from the Speech Emotion Recognition Case

ICASSP 2021accepted

With the rise of deep learning, deep knowledge transfer has emerged as one of the most effective techniques for getting state-of-the-art performance using deep neural networks. A lot of recent research has focused on understanding the mechanisms of transfer learning in the image and language domains…

Cited by 0SourceScholar
2020

Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition Models

ICASSP 2020accepted

The development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this…

Cited by 0SourceScholar
2020

Hierarchical Attention Transfer Networks for Depression Assessment from Speech

ICASSP 2020accepted

A growing area of mental health research is the search for speech-based objective markers for conditions such as depression. However, when combined with machine learning, this search can be challenging due to a limited amount of annotated training data. In this paper, we propose a novel crosstask ap…

Cited by 0SourceScholar
2020

Ordinal Learning for Emotion Recognition in Customer Service Calls

ICASSP 2020accepted

Approaches toward ordinal speech emotion recognition (SER) tasks are commonly based on the categorical classification algorithms, where the rank-order emotions are arbitrarily treated as independent categories. To employ the ordinal information between emotional ranks, we propose to model the ordina…

Cited by 0SourceScholar
2020

Stargan for Emotional Speech Conversion: Validated by Data Augmentation of End-To-End Emotion Recognition

ICASSP 2020accepted

In this paper, we propose an adversarial network implementation for speech emotion conversion as a data augmentation method, validated by a multi-class speech affect recognition task. In our setting, we do not assume the availability of parallel data, and we additionally make it a priority to exploi…

Cited by 0SourceScholar
2019

Attention-augmented End-to-end Multi-task Learning for Emotion Prediction from Speech

ICASSP 2019accepted

Despite the increasing research interest in end-to-end learning systems for speech emotion recognition, conventional systems either suffer from the overfitting due in part to the limited training data, or do not explicitly consider the different contributions of automatically learnt representations…

Cited by 0SourceScholar
2019

Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic Scenes

ICASSP 2019accepted

The goal of Acoustic Scene Classification (ASC) is to recognise the environment in which an audio waveform has been recorded. Recently, deep neural networks have been applied to ASC and have achieved state-of-the-art performance. However, few works have investigated how to visualise and understand w…

Cited by 0SourceScholar
2019

Context Modelling Using Hierarchical Attention Networks for Sentiment and Self-assessed Emotion Detection in Spoken Narratives

ICASSP 2019accepted

Automatic detection of sentiment and affect in personal narratives through word usage has the potential to assist in the automated detection of change in psychotherapy. Such a tool could, for instance, provide an efficient, objective measure of the time a person has been in a positive or negative st…

Cited by 0SourceScholar
2019

Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality

ICASSP 2019accepted

Despite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the trai…

Cited by 0SourceScholar
2018

End-to-End Speech Emotion Recognition Using Deep Neural Networks

ICASSP 2018accepted

Affect recognition is an important component towards the better interaction between human and machines. Applications of emotion recognition in speech can be found in several areas such as human computer interaction and call centres. In recent years, Deep Neural Networks (DNN) have been used with gre…

Cited by 0SourceScholar
2018

Multimodal Bag-of-Words for Cross Domains Sentiment Analysis

ICASSP 2018accepted

The advantages of using cross domain data when performing text-based sentiment analysis have been established; however, similar findings have yet to be observed when performing multimodal sentiment analysis. A potential reason for this is that systems based on feature extracted from speech and facia…

Cited by 46SourceScholar
2018

Towards Conditional Adversarial Training for Predicting Emotions from Speech

ICASSP 2018accepted

Motivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consi…

Cited by 0SourceScholar
2018

What is my Dog Trying to Tell Me? the Automatic Recognition of the Context and Perceived Emotion of Dog Barks

ICASSP 2018accepted

A wide range of research disciplines are deeply interested in the measurement of animal emotions, including evolutionary zoology, affective neuroscience and comparative psychology. However, only a few studies have investigated the effect of phenomena such as emotion on the acoustic parameters of (no…

Cited by 0SourceScholar
2017

Automatic multi-lingual arousal detection from voice applied to real product testing applications

ICASSP 2017accepted

A method is presented which applies Long Short-Term Memory Recurrent Neural Networks on real market-research voice recordings in order to automatically predict emotional arousal from speech. While most previous work has dealt with evaluations of algorithms within the same speech corpus, the novelty…

Cited by 6SourceScholar
2017

Multi-task deep neural network with shared hidden layers: Breaking down the wall between emotion representations

ICASSP 2017accepted

Emotion representations are psychological constructs for modelling, analysing, and recognising emotion, being one essential element of affect. Due to its complexity, the boundaries between different emotion concepts are often fuzzy, which is also reflected in the diversification of emotion databases…

Cited by 0SourceScholar
2017

Prediction-based learning for continuous emotion recognition in speech

ICASSP 2017accepted

In this paper, a prediction-based learning framework is proposed for a continuous prediction task of emotion recognition from speech, which is one of the key components of affective computing in multimedia. The main goal of this framework is to utmost exploit the individual advantages of different r…

Cited by 0SourceScholar
2017

Reconstruction-error-based learning for continuous emotion recognition in speech

ICASSP 2017accepted

To advance the performance of continuous emotion recognition from speech, we introduce a reconstruction-error-based (RE-based) learning framework with memory-enhanced Recurrent Neural Networks (RNN). In the framework, two successive RNN models are adopted, where the first model is used as an autoenc…

Cited by 0SourceScholar
2016

Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network

ICASSP 2016accepted

The automatic recognition of spontaneous emotions from speech is a challenging task. On the one hand, acoustic features need to be robust enough to capture the emotional content for various styles of speaking, and while on the other, machine learning algorithms need to be insensitive to outliers whi…

Cited by 0SourceScholar
2016

Cross lingual speech emotion recognition using canonical correlation analysis on principal component subspace

ICASSP 2016accepted

This paper proposes an analytical approach based on Kernel Canonical Correlation Analysis (KCCA) for domain adaptation. To generate paired instances for KCCA, we mapped source and target data onto both source and target principal components. We performed pair-wise domain adaptation between four emot…

Cited by 0SourceScholar
2016

Enhanced semi-supervised learning for multimodal emotion recognition

ICASSP 2016accepted

Semi-Supervised Learning (SSL) techniques have found many applications where labeled data is scarce and/or expensive to obtain. However, SSL suffers from various inherent limitations that limit its performance in practical applications. A central problem is that the low performance that a classifier…

Cited by 0SourceScholar
2016

Semi-autonomous data enrichment based on cross-task labelling of missing targets for holistic speech analysis

ICASSP 2016accepted

In this work, we propose a novel approach for large-scale data enrichment, with the aim to address a major shortcoming of current research in computational paralinguistics, namely, looking at speaker attributes in isolation although strong interdependencies between them exist. The scarcity of multi-…

Cited by 0SourceScholar
2016

Wavelet features for classification of vote snore sounds

ICASSP 2016accepted

Location and form of the upper airway obstruction is essential for a targeted therapy of obstructive sleep apnea (OSA). Utilizing snore sounds (SnS) to reveal the pathological characters of OSA patients has been the subject of scientific research for several decades. Fewer studies exist on the evalu…

Cited by 0SourceScholar
2015

A novel approach for automatic acoustic novelty detection using a denoising autoencoder with bidirectional LSTM neural networks

ICASSP 2015accepted

Acoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel unsupervised approach based on a denoising autoencoder. In our approach auditory spectral features are process…

Cited by 0SourceScholar