← Search

Hakan Erdogan

18 accepted papers

2024

Binaural Angular Separation Network

ICASSP 2024accepted

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omnidirectional microphones without needing to collect real RIRs. By relying…

Cited by 0SourceScholar
2024

Quantifying The Effect Of Simulator-Based Data Augmentation For Speech Recognition On Augmented Reality Glasses

ICASSP 2024accepted

Augmented reality (AR) glasses have an immense potential for enhancing conversations by leveraging speech recognition to display real-time transcription or translation, for example, to assist people with hearing impairments or for people conversing in a non-native language. For deployment in real en…

Cited by 6SourceScholar
2023

Guided Speech Enhancement Network

ICASSP 2023accepted

High quality speech capture has been widely studied for both voice communication and human computer interface reasons. To improve the capture performance, we can often find multi-microphone speech enhancement techniques deployed on various devices. Multi-microphone speech enhancement problem is ofte…

Cited by 0SourceScholar
2022

Adapting Speech Separation to Real-World Meetings using Mixture Invariant Training

ICASSP 2022accepted

The recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models because it does not require ground-truth isolated reference sources. In this paper, we investigate using MixIT to adapt a separation model on real far-field overlapp…

Cited by 26SourceScholar
2021

End-To-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

ICASSP 2021accepted

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforward handling of discriminative training, unlike traditional clustering-based diari…

Cited by 0SourceScholar
2021

Sound Event Detection and Separation: A Benchmark on Desed Synthetic Soundscapes

ICASSP 2021accepted

We propose a benchmark of state-of-the-art sound event detection systems (SED). We design synthetic evaluation sets to focus on specific sound event detection challenges. We analyze the performance of the submissions to DCASE 2020 Task 4 as a function of time-related modifications (time position of…

Cited by 0SourceScholar
2021

What's all the Fuss about Free Universal Sound Separation Data?

ICASSP 2021accepted

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio data drawn from 357 classes, which are used to create mixtur…

Cited by 0SourceScholar
2020

Performance Study of a Convolutional Time-Domain Audio Separation Network for Real-Time Speech Denoising

ICASSP 2020accepted

Time-domain audio separation networks based on dilated temporal convolutions have recently been shown to perform very well compared to methods that are based on a time-frequency representation in speech separation tasks, even outperforming an oracle binary time-frequency mask of the speakers. This p…

Cited by 0SourceScholar
2020

Unsupervised Sound Separation Using Mixture Invariant Training

NeurIPS 2020spotlight

In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from synthetic mixtures created by adding up isolated ground-truth sou…

2019

Low-latency Speaker-independent Continuous Speech Separation

ICASSP 2019accepted

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of which contains no overlapping speech segment. A separated, or cleaned, version of e…

Cited by 0SourceScholar
2019

Single-channel Speech Extraction Using Speaker Inventory and Attention Network

ICASSP 2019accepted

Neural network-based speech separation has received a surge of interest in recent years. Previously proposed methods either are speaker independent or extract a target speaker's voice by using his or her voice snippet. In applications such as home devices or office meeting transcriptions, a possible…

Cited by 76SourceScholar
2018

Exploring Practical Aspects of Neural Mask-Based Beamforming for Far-Field Speech Recognition

ICASSP 2018accepted

This work examines acoustic beamformers employing neural networks (NNs) for mask prediction as front -end for automatic speech recognition (ASR) systems for practical scenarios like voice-enabled home devices. To test the versatility of the mask predicting network, the system is evaluated with diffe…

Cited by 77SourceScholar
2018

Multi-Microphone Neural Speech Separation for Far-Field Multi-Talker Speech Recognition

ICASSP 2018accepted

This paper describes a neural network approach to far-field speech separation using multiple microphones. Our proposed approach is speaker-independent and can learn to implicitly figure out the number of speakers constituting an input speech mixture. This is realized by utilizing the permutation inv…

Cited by 0SourceScholar
2017

Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition

ICASSP 2017accepted

Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple microphones and performing beamforming over them. In this paper, we propose to use…

Cited by 0SourceScholar
2016

Deep beamforming networks for multi-channel speech recognition

ICASSP 2016accepted

Despite the significant progress in speech recognition enabled by deep neural networks, poor performance persists in some scenarios. In this work, we focus on far-field speech recognition which remains challenging due to high levels of noise and reverberation in the captured speech signals. We propo…

Cited by 0SourceScholar
2015

Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks

ICASSP 2015accepted

Separation of speech embedded in non-stationary interference is a challenging problem that has recently seen dramatic improvements using deep network-based methods. Previous work has shown that estimating a masking function to be applied to the noisy spectrum is a viable approach that can be improve…

Cited by 0SourceScholar