← Search

Stefano Squartini

5 accepted papers

2024

One Model to Rule Them All ? Towards End-to-End Joint Speaker Diarization and Speech Recognition

ICASSP 2024accepted

This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any number of speakers, effectively solving "who spoke what, whe…

Cited by 0SourceScholar
2023

Multi-Channel Speaker Extraction with Adversarial Training: The Wavlab Submission to The Clarity ICASSP 2023 Grand Challenge

ICASSP 2023accepted

In this work we detail our submission to the Clarity ICASSP 2023 grand challenge, in which participants have to develop a strong target speech enhancement system for hearing-aid (HA) devices in noisy-reverberant environments. Our system builds on our previous submission at the Second Clarity Enhance…

Cited by 0SourceScholar
2022

Learning Filterbanks for End-to-End Acoustic Beamforming

ICASSP 2022accepted

Recent work on monaural source separation has shown that performance can be increased by using fully learned filterbanks with short windows. On the other hand it is widely known that, for conventional beamforming techniques, performance increases with long analysis windows. This applies also to most…

Cited by 0SourceScholar
2019

End-to-end Binaural Sound Localisation from the Raw Waveform

ICASSP 2019accepted

A novel end-to-end binaural sound localisation approach is proposed which estimates the azimuth of a sound source directly from the waveform. Instead of employing hand-crafted features commonly employed for binaural sound localisation, such as the interaural time and level difference, our end-to-end…

Cited by 63SourceScholar
2015

A novel approach for automatic acoustic novelty detection using a denoising autoencoder with bidirectional LSTM neural networks

ICASSP 2015accepted

Acoustic novelty detection aims at identifying abnormal/novel acoustic signals which differ from the reference/normal data that the system was trained with. In this paper we present a novel unsupervised approach based on a denoising autoencoder. In our approach auditory spectral features are process…

Cited by 0SourceScholar