← Search

Tim Fingscheidt

25 accepted papers

2026

ICASSP 2026 URGENT Speech Enhancement Challenge

ICASSP 2026poster

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task definitions, datasets, baseline systems, evaluation protocols, an…

Cited by 0SourcePDFScholar
2026

Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations

ICLR 2026poster

Large language models (LLMs) demand extensive memory capacity during both fine-tuning and inference. To enable memory-efficient fine-tuning, existing methods apply block-wise quantization techniques, such as NF4 and AF4, to the network weights. We show that these quantization techniques incur subopt…

Cited by 3SourcecodeScholar
2024

Distributed Semantic Segmentation with Efficient Joint Source and Task Decoding

ECCV 2024poster

"Distributed computing in the context of deep neural networks (DNNs) implies the execution of one part of the network on edge devices and the other part typically on a large-scale cloud platform. Conventional methods propose to employ a serial concatenation of a learned image and source encoder, the…

2024

Efficient High-Performance Bark-Scale Neural Network for Residual Echo and Noise Suppression

ICASSP 2024accepted

In recent years, the introduction of neural networks (NNs) into the field of speech enhancement has brought significant improvements. However, many of the proposed methods are quite demanding in terms of computational complexity and memory footprint. For the application in dedicated communication de…

Cited by 0SourceScholar
2022

Deep Residual Echo Suppression and Noise Reduction: A Multi-Input FCRN Approach in a Hybrid Speech Enhancement System

ICASSP 2022accepted

Deep neural network (DNN)-based approaches to acoustic echo cancellation (AEC) and hybrid speech enhancement systems have gained increasing attention recently, introducing significant performance improvements to this research field. Using the fully convolutional recurrent network (FCRN) architecture…

Cited by 0SourceScholar
2022

Detecting Adversarial Perturbations in Multi-Task Perception

IROS 2022poster

While deep neural networks (DNNs) achieve impressive performance on environment perception tasks, their sensitivity to adversarial perturbations limits their use in practical applications. In this paper, we (i) propose a novel adversarial perturbation detection scheme based on multi-task perception…

Cited by 23SourcecodeScholar
2021

A New DCASE 2017 Rare Sound Event Detection Benchmark Under Equal Training Data: CRNN With Multi-Width Kernels

ICASSP 2021accepted

Rare sound event detection (rare SED) deals with obtaining valuable information from data consisting mostly of acoustic background noises. It has meanwhile a long research history and was part of the DCASE 2017 Challenge. State-of-the-art performance is currently reached using a stacked combination…

Cited by 0SourceScholar
2021

AEC in A Netshell: on Target and Topology Choices for FCRN Acoustic Echo Cancellation

ICASSP 2021accepted

Acoustic echo cancellation (AEC) algorithms have a long-term steady role in signal processing, with approaches improving the performance of applications such as automotive hands-free systems, smart home and loudspeaker devices, or web conference systems. Just recently, very first deep neural network…

Cited by 0SourceScholar
2020

A Multichannel Kalman-Based Wiener Filter Approach for Speaker Interference Reduction in Meetings

ICASSP 2020accepted

Recording a meeting and obtaining clean speech signals of each speaker is a challenging task. Even with a multichannel recording, in which all speakers are equipped with a close-talk microphone, speech of an active speaker still couples not only into his dedicated microphone, but also into all other…

Cited by 0SourceScholar
2020

Beyond the Dcase 2017 Challenge on Rare Sound Event Detection: A Proposal for a More Realistic Training and Test Framework

ICASSP 2020accepted

There are many ways to evaluate rare sound event detection (SED) approaches, e.g., the DCASE 2017 challenge provides a widely employed framework. This paper proposes a rare SED training and test framework, which is reflecting an SED application in a more realistic way. Our setup gets rid of too much…

Cited by 0SourceScholar
2020

Fully Convolutional Recurrent Networks for Speech Enhancement

ICASSP 2020accepted

Convolutional recurrent neural networks (CRNs) using convolutional encoder-decoder (CED) structures have shown promising performance for single-channel speech enhancement. These CRNs handle temporal modeling through integrating long short-term memory (LSTM) layers in between convolutional encoder an…

Cited by 0SourceScholar
2020

Self-Supervised Monocular Depth Estimation: Solving the Dynamic Object Problem by Semantic Guidance

ECCV 2020poster

Self-supervised monocular depth estimation presents a powerful method to obtain 3D scene information from single camera images, which is trainable on arbitrary image sequences without requiring depth labels, e.g., from a LiDAR sensor. In this work we present a new self-supervised semantically-guided…

2019

Improved Measurement Noise Covariance Estimation for N-channel Feedback Cancellation Based on the Frequency Domain Adaptive Kalman Filter

ICASSP 2019accepted

Acoustic feedback cancellation has gained a major and steady role in the research fields of signal processing over the past decades, since it is inevitable for numerous applications such as hearing aids or in-car communication systems. In this paper, we investigate measurement noise covariance estim…

Cited by 0SourceScholar
2019

Learning to Dequantize Speech Signals by Primal-dual Networks: an Approach for Acoustic Sensor Networks

ICASSP 2019accepted

We introduce a method to improve the quality of simple scalar quantization in the context of acoustic sensor networks by combining ideas from sparse reconstruction, artificial neural networks and weighting filters. We start from the observation that optimization methods based on sparse reconstructio…

Cited by 0SourceScholar
2018

A Priori SNR Estimation Using Discriminative Non-Negative Matrix Factorization

ICASSP 2018accepted

A priori signal-to-noise ratio (SNR) contains critical information about the single-channel mixture of a speech and noise signal, and can be used by speech enhancement algorithms. In this paper, we propose a novel a priori SNR estimator using the estimates obtained from discriminative non-negative m…

Cited by 0SourceScholar
2018

A Simple Cepstral Domain DNN Approach to Artificial Speech Bandwidth Extension

ICASSP 2018accepted

In this work, we present a simple deep neural network (DNN)-based regression approach to artificial speech bandwidth extension (ABE) in the frequency domain for estimating missing speech components in the range 4 ... 7 kHz. The upper band (UB) spectral magnitudes are found by first estimating the UB…

Cited by 0SourceScholar
2018

An Efficient Residual Echo Suppression for Multi-Channel Acoustic Echo Cancellation Based on the Frequency-Domain Adaptive Kalman Filter

ICASSP 2018accepted

Emerging use cases, such as keyword spotting while listening to FM radio, or participating in a teleconference utilizing the hands-free system in a vehicle, require the utilization of multi-channel acoustic echo cancellation (AEC). Addressing the typically remaining residual echo, it is common pract…

Cited by 0SourceScholar
2016

A subjective listening test of six different artificial bandwidth extension approaches in English, Chinese, German, and Korean

ICASSP 2016accepted

In studies on artificial bandwidth extension (ABE), there is a lack of international coordination in subjective tests between multiple methods and languages. Here we present the design of absolute category rating listening tests evaluating 12 ABE variants of six approaches in multiple languages, nam…

Cited by 21SourceScholar
2016

Evaluating instrumental measures of speech quality using Bayesian model selection: Correlations can be misleading!

ICASSP 2016accepted

Choosing among competing models of collected data is crucial for all sciences. In the last decade there has been an increasing tendency to use Bayesian methods throughout many fields. When assessing the performance of instrumental measures of speech quality, classical measures such as correlation co…

Cited by 0SourceScholar
2016

Soft linear discriminant analysis (SLDA) for pattern recognition with ambiguous reference labels: Application to social signal processing

ICASSP 2016accepted

While most pattern recognition approaches are designed and trained with clearly defined reference labels, there are a few new applications working with ambiguous ones. Since the linear discriminant analysis (LDA) is one of the most utilized methods in pattern recognition to reduce the dimensionality…

Cited by 0SourceScholar
2016

System-compatible robustness improvement for new generation dect decoders by G.722 soft-decision decoding

ICASSP 2016accepted

The ITU-T Recommendation G.722 about subband adaptive differential pulse code modulation (SB-ADPCM) is the mandatory wideband speech codec in the new generation digital enhanced cordless telephony (NG-DECT). Although in ADPCM the difference signal instead of the original signal is quantized and adap…

Cited by 0SourceScholar
2015

Acoustic event source localization for surveillance in reverberant environments supported by an event onset detection

ICASSP 2015accepted

This contribution presents a robust approach to acoustic event source localization for surveillance under reverberant environmental conditions. In particular, we support the classical generalized cross-correlation algorithm with phase transform weighting (GCC-PHAT) and the steered response power (SR…

Cited by 0SourceScholar