← Search

Najim Dehak

29 accepted papers

2025

Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings

ICASSP 2025accepted

In this study, we explored the use of spectrograms to represent handwriting signals for assessing neurodegenerative diseases, including 42 healthy controls (CTL), 35 subjects with Parkinson’s Disease (PD), 21 with Alzheimer’s Disease (AD), and 15 with Parkinson’s Disease Mimics (PDM). We applied CNN…

Cited by 0SourceScholar
2025

Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation

EMNLP 2025

We present Paired by the Teacher (PbT), a two-stage teacher–student pipeline that synthesizes accurate input–output pairs without human labels or parallel data. In many low-resource natural language generation (NLG) scenarios, practitioners may have only raw outputs, like highlights, recaps, or ques

Cited by 0SourcePDFScholar
2025

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

ICASSP 2025accepted

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the g…

Cited by 0SourceScholar
2025

SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

ICASSP 2025accepted

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both a…

Cited by 0SourceScholar
2025

Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More

ICASSP 2025accepted

With the recent advancements in speech recognition, it is crucial to ensure these systems are free from performance biases against any speaker subgroups. This study examined the performance of twenty variants of seven Automatic Speech Recognition models across four datasets in English language: L2 A…

Cited by 0SourceScholar
2024

CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing

NeurIPS 2024poster

We introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to standard fine-tuning methods that optimize for downstream models, CA-SSLR integrates language and speaker embeddings from…

Cited by 0SourcePDFScholar
2024

DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction

ICASSP 2024accepted

Common target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative…

Cited by 0SourceScholar
2024

Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation Pipeline

COLING 2024main

The growing emphasis on fairness in speech-processing tasks requires datasets with speakers from diverse subgroups that allow training and evaluating fair speech technology systems. However, creating such datasets through manual annotation can be costly. To address this challenge, we present a semi-…

2021

Align or attend? Toward More Efficient and Accurate Spoken Word Discovery Using Speech-to-Image Retrieval

ICASSP 2021accepted

Multimodal word discovery (MWD) is often treated as a byproduct of the speech-to-image retrieval problem. However, our theoretical analysis shows that some kind of alignment/attention mechanism is crucial for a MWD system to learn meaningful word-level representation. We verify our theory by conduct…

Cited by 0SourceScholar
2021

CopyPaste: An Augmentation Method for Speech Emotion Recognition

ICASSP 2021accepted

Data augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expensive and challenging. This study proposes CopyPaste, a perceptually motivated no…

Cited by 0SourceScholar
2021

Focus on the Present: A Regularization Method for the ASR Source-Target Attention Layer

ICASSP 2021accepted

This paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classification (CTC) and attention training. Our method is based on the fact that both, CTC and source-target attention, are acting…

Cited by 0SourceScholar
2021

How Phonotactics Affect Multilingual and Zero-Shot ASR Performance

ICASSP 2021accepted

The idea of combining multiple languages’ recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been shown to leverage multilingual data well in IPA transcription…

Cited by 0SourceScholar
2021

Improving Reconstruction Loss Based Speaker Embedding in Unsupervised and Semi-Supervised Scenarios

ICASSP 2021accepted

Text-to-speech (TTS) models trained to minimize the spectrogram reconstruction loss can learn speaker embeddings without explicit speaker identity supervision, unlike x-vector speaker identification (SID) systems. Leveraging this way of speaker embedding learning can be useful in unsupervised or sem…

Cited by 0SourceScholar
2021

Perceptual Loss Based Speech Denoising with an Ensemble of Audio Pattern Recognition and Self-Supervised Models

ICASSP 2021accepted

Deep learning based speech denoising still suffers from the challenge of improving perceptual quality of enhanced signals. We introduce a generalized framework called Perceptual Ensemble Regularization Loss (PERL) built on the idea of perceptual losses. Perceptual loss discourages distortion to cert…

Cited by 0SourceScholar
2020

Feature Enhancement with Deep Feature Losses for Speaker Verification

ICASSP 2020accepted

Speaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which opt…

Cited by 0SourceScholar
2020

Unsupervised Feature Enhancement for Speaker Verification

ICASSP 2020accepted

The task of making speaker verification systems robust to adverse scenarios remains a challenging and an active area of research. We developed an unsupervised feature enhancement approach in log-filter bank space with the end goal of improving speaker verification performance. We experimented with u…

Cited by 0SourceScholar
2020

Using X-Vectors to Automatically Detect Parkinson's Disease from Speech

ICASSP 2020accepted

The promise of new neuroprotective treatments to stop or slow the advance of Parkinson's Disease (PD) urges for new biomarkers or detection schemes that can deliver a faster diagnosis. Given that speech is affected by PD, the combination of deep neural networks and speech processing can provide auto…

Cited by 0SourceScholar
2020

X-Vectors Meet Emotions: A Study On Dependencies Between Emotion and Speaker Recognition

ICASSP 2020accepted

In this work, we explore the dependencies between speaker recognition and emotion recognition. We first show that knowledge learned for speaker recognition can be reused for emotion recognition through transfer learning. Then, we show the effect of emotion on speaker recognition. For emotion recogni…

Cited by 0SourceScholar
2019

Attentive Filtering Networks for Audio Replay Attack Detection

ICASSP 2019accepted

An attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The ASVspoof 2017 Challenge focused specifically on replay attacks, with the intenti…

Cited by 0SourceScholar
2019

Cycle-GANs for Domain Adaptation of Acoustic Features for Speaker Recognition

ICASSP 2019accepted

It is well known that domain mismatch between the training and evaluation data hinders the performance of any machine learning system. Various factors contribute to domain mismatch. In speaker recognition systems, it mainly occurs due to the mismatch in recording conditions and language. Most speake…

Cited by 0SourceScholar
2019

Investigation on Neural Bandwidth Extension of Telephone Speech for Improved Speaker Recognition

ICASSP 2019accepted

We extend our previous work on training mixed-bandwidth (BW) speaker recognition system by predicting missing information in upperband (UB) of upsampled telephone speech. Mixed-BW systems combine speech from narrowband (NB) and wideband (WB) speech corpora by basic upsampling of NB speech with low-p…

Cited by 0SourceScholar
2019

Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition

ICASSP 2019accepted

In this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained language model (LM). Our proposed fusion methods focus on the memory cell state and the hidden state in the seq2seq decoder long short-term memory (LSTM), and the memory cell state is updated by the LM…

Cited by 0SourceScholar
2018

Characterizing Performance of Speaker Diarization Systems on Far-Field Speech Using Standard Methods

ICASSP 2018accepted

To date, the bulk of research on speaker diarization has been conducted on telephone or near-field speech. As the need for technologies capable of handling conversational speech increases, it is necessary to establish the performance of state-of-the-art systems in this domain. In this work we evalua…

Cited by 0SourceScholar
2018

Joint Verification-Identification in end-to-end Multi-Scale CNN Framework for Topic Identification

ICASSP 2018accepted

We present an end-to-end multi-scale Convolutional Neural Network (CNN) framework for topic identification (topic ID). In this work, we examined multi -scale CNN for classification using raw text input. Topical word embeddings are learnt at multiple scales using parallel convolutional layers. A tech…

Cited by 0SourceScholar
2018

Measuring Uncertainty in Deep Regression Models: The Case of Age Estimation from Speech

ICASSP 2018accepted

Age estimation from speech recently received a lot of attention. Approaches such as i-vectors and deep learning have been successfully applied to this task achieving great performance. However, one drawback of those methods is that they produce a hard age estimation without any kind of confidence me…

Cited by 1SourceScholar
2017

An empirical evaluation of zero resource acoustic unit discovery

ICASSP 2017accepted

Acoustic unit discovery (AUD) is a process of automatically identifying a categorical acoustic unit inventory from speech and producing corresponding acoustic unit tokenizations. AUD provides an important avenue for unsupervised acoustic model training in a zero resource setting where expert-provide…

Cited by 0SourceScholar
2017

Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease

ICASSP 2017accepted

Information from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized…

Cited by 35SourceScholar
2017

Topic identification of spoken documents using unsupervised acoustic unit discovery

ICASSP 2017accepted

This paper investigates the application of unsupervised acoustic unit discovery for topic identification (topic ID) of spoken audio documents. The acoustic unit discovery method is based on a non-parametric Bayesian phone-loop model that segments a speech utterance into phone-like categories. The di…

Cited by 0SourceScholar