← Search

Konstantinos Drossos

7 accepted papers

2026

BEYOND OMNIDIRECTIONAL: NEURAL AMBISONICS ENCODING FOR ARBITRARY MICROPHONE DIRECTIVITY PATTERNS USING CROSS-ATTENTION

ICASSP 2026oral

We present a deep neural network approach for encoding microphone array signals into Ambisonics that generalizes to arbitrary microphone array configurations with fixed microphone count but varying locations and frequency-dependent directional characteristics. Unlike previous methods that rely only…

Cited by 0SourcePDFScholar
2025

Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays

ICASSP 2025accepted

Using deep neural networks (DNNs) for encoding of microphone array (MA) signals to the Ambisonics spatial audio format can surpass certain limitations of established conventional methods, but existing DNN-based methods need to be trained separately for each MA. This paper proposes a DNN-based method…

Cited by 0SourceScholar
2022

Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual Phrases

ICASSP 2022accepted

We investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unaligned and unannotated audio clips and their captions by scoring the similarities b…

Cited by 0SourceScholar
2021

Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags

ICASSP 2021accepted

Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags associated with audio do not employ text processing models that ar…

Cited by 0SourceScholar
2020

Sound Event Detection Via Dilated Convolutional Recurrent Neural Networks

ICASSP 2020accepted

Convolutional recurrent neural networks (CRNNs) have achieved state-of-the-art performance for sound event detection (SED). In this paper, we propose to use a dilated CRNN, namely a CRNN with a dilated convolutional kernel, as the classifier for the task of SED. We investigate the effectiveness of d…

Cited by 0SourceScholar
2018

Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency Mask

ICASSP 2018accepted

Singing voice separation based on deep learning relies on the usage of time-frequency masking. In many cases the masking process is not a learnable function or is not encapsulated into the deep learning optimization. Consequently, most of the existing methods rely on a post processing step using the…

Cited by 0SourceScholar