← Search

Hassan Taherian

8 accepted papers

2025

Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition

ICASSP 2025accepted

Despite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating individual talkers but, as a frontend, it introduces processing arti…

Cited by 0SourceScholar
2025

Robust Frame-level Speaker Localization in Reverberant and Noisy Environments by Exploiting Phase Difference Losses

ICASSP 2025accepted

This paper investigates robust speaker localization at the frame level on the basis of complex spectral mapping, which is capable of learning both the magnitude and phase of the target signal. Unlike prevailing deep learning methods for speaker localization, we perform MIMO (multi-input multi-output…

Cited by 0SourceScholar
2024

Leveraging Sound Localization to Improve Continuous Speaker Separation

ICASSP 2024accepted

Continuous speaker separation aims to separate overlapping speakers in real-world environments like meetings, but it often falls short in isolating speech segments of a single speaker. This leads to split signals that adversely affect downstream applications such as automatic speech recognition and…

Cited by 9SourceScholar
2023

Breaking the Trade-Off in Personalized Speech Enhancement With Cross-Task Knowledge Distillation

ICASSP 2023accepted

Personalized speech enhancement (PSE) models achieve promising results compared with unconditional speech enhancement models due to their ability to remove interfering speech in addition to background noise. Unlike unconditional speech enhancement, causal PSE models may occasionally remove the targe…

Cited by 0SourceScholar
2023

Multi-Resolution Location-Based Training for Multi-Channel Continuous Speech Separation

ICASSP 2023accepted

The performance of automatic speech recognition (ASR) systems severely degrades when multi-talker speech overlap occurs. In meeting environments, speech separation is typically performed to improve the robustness of ASR systems. Recently, location-based training (LBT) was proposed as a new training…

Cited by 8SourceScholar
2022

Location-Based Training for Multi-Channel Talker-Independent Speaker Separation

ICASSP 2022accepted

Permutation-invariant training (PIT) is a dominant approach for addressing the permutation ambiguity problem in talker-independent speaker separation. Leveraging spatial information afforded by microphone arrays, we propose a new training approach to resolving permutation ambiguities for multi-chann…

Cited by 0SourceScholar
2022

One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech Enhancement

ICASSP 2022accepted

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect with friends and families. Single-channel personalized speech enhancement (PSE) methods show promising results compared…

Cited by 0SourceScholar
2021

Time-Domain Loss Modulation Based on Overlap Ratio for Monaural Conversational Speaker Separation

ICASSP 2021accepted

Existing speaker separation methods deliver excellent performance on fully overlapped signal mixtures. To apply these methods in daily conversations that include occasional concurrent speakers, recent studies incorporate both overlapped and non-overlapped segments in the training data. However, such…

Cited by 0SourceScholar