← Search

Tuomas Virtanen

25 accepted papers

2026

BEYOND OMNIDIRECTIONAL: NEURAL AMBISONICS ENCODING FOR ARBITRARY MICROPHONE DIRECTIVITY PATTERNS USING CROSS-ATTENTION

ICASSP 2026oral

We present a deep neural network approach for encoding microphone array signals into Ambisonics that generalizes to arbitrary microphone array configurations with fixed microphone count but varying locations and frequency-dependent directional characteristics. Unlike previous methods that rely only…

Cited by 0SourcePDFScholar
2025

A decade of DCASE: Achievements, practices, evaluations and future challenges

ICASSP 2025accepted

This paper introduces briefly the history and growth of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, workshop, research area and research community. Created in 2013 as a data evaluation challenge, DCASE has become a major research topic in the Audio and Acoustic…

Cited by 0SourceScholar
2025

Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays

ICASSP 2025accepted

Using deep neural networks (DNNs) for encoding of microphone array (MA) signals to the Ambisonics spatial audio format can surpass certain limitations of established conventional methods, but existing DNN-based methods need to be trained separately for each MA. This paper proposes a DNN-based method…

Cited by 0SourceScholar
2024

Attention-Driven Multichannel Speech Enhancement in Moving Sound Source Scenarios

ICASSP 2024accepted

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering techniques designed for dynamic settings. Specifically, we stu…

Cited by 0SourceScholar
2024

Neural Ambisonics Encoding For Compact Irregular Microphone Arrays

ICASSP 2024accepted

Ambisonics encoding of microphone array signals can enable various spatial audio applications, such as virtual reality or telepresence, but it is typically designed for uniformly-spaced spherical microphone arrays. This paper proposes a method for Ambisonics encoding that uses a deep neural network…

Cited by 12SourceScholar
2023

STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events

NeurIPS 2023poster

While direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker. This paper proposes an audio-visual sou…

2022

Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual Phrases

ICASSP 2022accepted

We investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unaligned and unannotated audio clips and their captions by scoring the similarities b…

Cited by 0SourceScholar
2021

A Curated Dataset of Urban Scenes for Audio-Visual Scene Analysis

ICASSP 2021accepted

This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multiple European cities, using the same equipment, in multiple locations for each scene, and is openly available. We also pres…

Cited by 0SourceScholar
2021

Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags

ICASSP 2021accepted

Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags associated with audio do not employ text processing models that ar…

Cited by 0SourceScholar
2021

Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic Projections

ICASSP 2021accepted

In this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Zero-shot learning in audio classification refers to classification problems that aim at recognizing audio instances of so…

Cited by 0SourceScholar
2020

Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks

ICASSP 2020accepted

Convolutional recurrent neural networks (CRNNs) have achieved state-of-the-art performance for sound event detection (SED). In this paper, we propose to use a dilated CRNN, namely a CRNN with a dilated convolutional kernel, as the classifier for the task of SED. We investigate the effectiveness of d…

Cited by 54SourceScholar
2019

Sound Event Envelope Estimation in Polyphonic Mixtures

ICASSP 2019accepted

Sound event detection is the task of identifying automatically the presence and temporal boundaries of sound events within an input audio stream. In the last years, deep learning methods have established themselves as the state-of-the-art approach for the task, using binary indicators during trainin…

Cited by 0SourceScholar
2018

Estimation of Time-Varying Room Impulse Responses of Multiple Sound Sources from Observed Mixture and Isolated Source Signals

ICASSP 2018accepted

This paper proposes a method for online estimation of time-varying room impulse responses (RIR) between multiple isolated sound sources and a far-field mixture. The algorithm is formulated as adaptive convolutive filtering in short-time Fourier transform (STFT) domain. We use the recursive least squ…

Cited by 0SourceScholar
2018

Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency Mask

ICASSP 2018accepted

Singing voice separation based on deep learning relies on the usage of time-frequency masking. In many cases the masking process is not a learnable function or is not encapsulated into the deep learning optimization. Consequently, most of the existing methods rely on a post processing step using the…

Cited by 0SourceScholar
2017

Active learning for sound event classification by clustering unlabeled data

ICASSP 2017accepted

This paper proposes a novel active learning method to save annotation effort when preparing material to train sound event classifiers. K-medoids clustering is performed on unlabeled sound segments, and medoids of clusters are presented to annotators for labeling. The annotated label for a medoid is…

Cited by 0SourceScholar
2017

Sound event detection using spatial features and convolutional recurrent neural network

ICASSP 2017accepted

This paper proposes to use low-level spatial features extracted from multichannel audio for sound event detection. We extend the convolutional recurrent neural network to handle more than one type of these multichannel features by learning from each of them separately in the initial stages. We show…

Cited by 0SourceScholar
2016

Recurrent neural networks for polyphonic sound event detection in real life recordings

ICASSP 2016accepted

In this paper we present an approach to polyphonic sound event detection in real life recordings based on bi-directional long short term memory (BLSTM) recurrent neural networks (RNNs). A single multilabel BLSTM RNN is trained to map acoustic features of a mixture signal consisting of sounds from mu…

Cited by 0SourceScholar
2015

Exemplar-based speech enhancement for deep neural network based automatic speech recognition

ICASSP 2015accepted

Deep neural network (DNN) based acoustic modelling has been successfully used for a variety of automatic speech recognition (ASR) tasks, thanks to its ability to learn higher-level information using multiple hidden layers. This paper investigates the recently proposed exemplar-based speech enhanceme…

Cited by 0SourceScholar
2015

Low-latency sound-source-separation using non-negative matrix factorisation with coupled analysis and synthesis dictionaries

ICASSP 2015accepted

For real-time or close to real-time applications, sound source separation can be performed on-line, where new frames of incoming data for a mixture signal are processed as they arrive, at very low delay. We propose an approach which generates the separation filters for short synthesis frames to achi…

Cited by 0SourceScholar
2015

Similarity induced group sparsity for non-negative matrix factorisation

ICASSP 2015accepted

Non-negative matrix factorisations are used in several branches of signal processing and data analysis for separation and classification. Sparsity constraints are commonly set on the model to promote discovery of a small number of dominant patterns. In group sparse models, atoms considered to belong…

Cited by 0SourceScholar
2015

Sound event detection in real life recordings using coupled matrix factorization of spectral representations and class activity annotations

ICASSP 2015accepted

Methods for detection of overlapping sound events in audio involve matrix factorization approaches, often assigning separated components to event classes. We present a method that bypasses the supervised construction of class models. The method learns the components as a non-negative dictionary in a…

Cited by 0SourceScholar