← Search

Vahid Ahmadi Kalkhorani

2 accepted papers

2025

Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition

ICASSP 2025accepted

Despite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating individual talkers but, as a frontend, it introduces processing arti…

Cited by 0SourceScholar
2024

Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain

ICASSP 2024accepted

We introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual…

Cited by 0SourceScholar