← Search

Hannes Gamper

25 accepted papers

2025

Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

AAAI 2025technical

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage i…

2025

Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment

ICASSP 2025accepted

In this paper, we investigate distillation and pruning methods to reduce model size for non-intrusive speech quality assessment based on self-supervised representations. Our experiments build on XLS-R-SQA, a speech quality assessment model using wav2vec 2.0 XLS-R embeddings. We retrain this model on…

Cited by 0SourceScholar
2025

Make Some Noise: Towards LLM audio reasoning and generation using sound tokens

ICASSP 2025accepted

Integrating audio comprehension and generation into large language models (LLMs) remains challenging due to the continuous nature of audio and the resulting high sampling rates. Here, we introduce a novel approach that combines Variational Quantization with Conditional Flow Matching to convert audio…

Cited by 0SourceScholar
2024

Adapting Frechet Audio Distance for Generative Music Evaluation

ICASSP 2024accepted

The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual quality is understudied. We show that FAD performance may b…

Cited by 0SourceScholar
2024

An Inverse Kinematics Algorithm With Smooth Task Switching for Redundant Robots

RA-L 2024

This paper presents an inverse kinematics approach that combines two well-known Jacobian based methods, the task-priority framework and an optimization-based approach, such that tracking and optimization tasks can be executed simultaneously. The novelty of the proposed algorithm lies in the ability

Cited by 6SourceScholar
2022

ICASSP 2022 Acoustic Echo Cancellation Challenge

ICASSP 2022accepted

The ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC challenge and it is enhanced by including mobile scenarios,…

Cited by 0SourceScholar
2022

Icassp 2022 Deep Noise Suppression Challenge

ICASSP 2022accepted

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-sourc…

Cited by 0SourceScholar
2021

Decoding Music Attention from "EEG Headphones": A User-Friendly Auditory Brain-Computer Interface

ICASSP 2021accepted

People enjoy listening to music as part of their life. This makes music an excellent choice for designing a user-friendly brain-computer interface (BCI) for long-term use. We propose a novel BCI system using music stimuli that relies on brain signals collected via Smartfones, an EEG recording device…

Cited by 0SourceScholar
2021

ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results

ICASSP 2021accepted

The ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performa…

Cited by 0SourceScholar
2021

ICASSP 2021 Deep Noise Suppression Challenge

ICASSP 2021accepted

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to tr…

Cited by 0SourceScholar
2021

Towards Efficient Models for Real-Time Deep Noise Suppression

ICASSP 2021accepted

With recent research advancements, deep learning models are be-coming attractive and powerful choices for speech enhancement in real-time applications. While state-of-the-art models can achieve outstanding results in terms of speech quality and background noise reduction, the main challenge is to ob…

Cited by 0SourceScholar
2020

Fast Acoustic Scattering Using Convolutional Neural Networks

ICASSP 2020accepted

Diffracted scattering and occlusion are important acoustic effects in interactive auralization and noise control applications, typically requiring expensive numerical simulation. We propose training a convolutional neural network to map from a convex scatterer's cross-section to a 2D slice of the re…

Cited by 0SourceScholar
2020

Predicting Word Error Rate for Reverberant Speech

ICASSP 2020accepted

Reverberation negatively impacts the performance of automatic speech recognition (ASR). Prior work on quantifying the effect of reverberation has shown that clarity (C50), a parameter that can be estimated from the acoustic impulse response, is correlated with ASR performance. In this paper we propo…

Cited by 0SourceScholar
2019

A Sparsity Measure for Echo Density Growth in General Environments

ICASSP 2019accepted

We study the detailed temporal evolution of echo density in impulse responses for applications in acoustic analysis and rendering on general environments. For this purpose, we propose a smooth sorted density measure that yields an intuitive trend of echo density growth with time. This is fitted with…

Cited by 0SourceScholar
2019

Blind Room Volume Estimation from Single-channel Noisy Speech

ICASSP 2019accepted

Recent work on acoustic parameter estimation indicates that geometric room volume can be useful for modeling the character of an acoustic environment. However, estimating volume from audio signals remains a challenging problem. Here we propose using a convolutional neural network model to estimate t…

Cited by 0SourceScholar
2019

Improving Binaural Ambisonics Decoding by Spherical Harmonics Domain Tapering and Coloration Compensation

ICASSP 2019accepted

A powerful and flexible approach to record or encode a spatial sound scene is through spherical harmonics (SHs), or Ambisonics. An SH-encoded scene can be rendered binaurally by applying SH-encoded head-related transfer functions (HRTFs). Limitations of the recording equipment or computational const…

Cited by 0SourceScholar
2019

Non-intrusive Speech Quality Assessment Using Neural Networks

ICASSP 2019accepted

Estimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this work, we present an investigation of the applicability of neura…

Cited by 0SourceScholar
2016

Applications of 3D spherical transforms to personalization of head-related transfer functions

ICASSP 2016accepted

Head-related transfer functions (HRTFs) depend on the shape of the human head and ears, motivating HRTF personalization methods that detect and exploit morphological similarities between subjects in an HRTF database and a new user. Prior work determined similarity from sets of morphological paramete…

Cited by 0SourceScholar
2016

BFGUI: An interactive tool for the synthesis and analysis of microphone array beamformers

ICASSP 2016accepted

Microphone arrays are beneficial for distant speech capture because the signals they capture can be exploited with beamforming to suppress noise and reverberation. The theory for the design and analysis of microphone arrays is well established, however the performance of a microphone array beamforme…

Cited by 0SourceScholar
2015

Dereverberation sweet spot dilation with combined channel equalization and beamforming

ICASSP 2015accepted

Beamforming and channel equalizers can be formulated as optimal multichannel filter-and-sum operations with different objective criteria. It has been shown in previous studies that the combination of both concepts under a common framework can yield results that combine both the spatial robustness of…

Cited by 0SourceScholar
2015

Estimation of multipath propagation delays and interaural time differences from 3-D head scans

ICASSP 2015accepted

The estimation of acoustic propagation delays from a sound source to a listener's ear entrances is useful for understanding and visualising the wave propagation along the surface of the head, and necessary for individualised spatial sound rendering. The interaural time difference (ITD) is of particu…

Cited by 0SourceScholar