← Search

Phani Sankar Nidadavolu

7 accepted papers

2024

Hot-Fixing Wake Word Recognition for End-to-End ASR Via Neural Model Reprogramming

ICASSP 2024accepted

This paper proposes two novel variants of neural reprogramming to enhance wake word recognition in streaming end-to-end ASR models without updating model weights. The first, "trigger-frame reprogramming", prepends the input speech feature sequence with the learned trigger-frames of the target wake w…

Cited by 0SourceScholar
2020

Feature Enhancement with Deep Feature Losses for Speaker Verification

ICASSP 2020accepted

Speaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which opt…

Cited by 0SourceScholar
2020

Unsupervised Feature Enhancement for Speaker Verification

ICASSP 2020accepted

The task of making speaker verification systems robust to adverse scenarios remains a challenging and an active area of research. We developed an unsupervised feature enhancement approach in log-filter bank space with the end goal of improving speaker verification performance. We experimented with u…

Cited by 0SourceScholar
2019

Cycle-GANs for Domain Adaptation of Acoustic Features for Speaker Recognition

ICASSP 2019accepted

It is well known that domain mismatch between the training and evaluation data hinders the performance of any machine learning system. Various factors contribute to domain mismatch. In speaker recognition systems, it mainly occurs due to the mismatch in recording conditions and language. Most speake…

Cited by 0SourceScholar
2019

Investigation on Neural Bandwidth Extension of Telephone Speech for Improved Speaker Recognition

ICASSP 2019accepted

We extend our previous work on training mixed-bandwidth (BW) speaker recognition system by predicting missing information in upperband (UB) of upsampled telephone speech. Mixed-BW systems combine speech from narrowband (NB) and wideband (WB) speech corpora by basic upsampling of NB speech with low-p…

Cited by 0SourceScholar
2017

Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease

ICASSP 2017accepted

Information from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized…

Cited by 0SourceScholar
2017

On the impact of non-modal phonation on phonological features

ICASSP 2017accepted

Different modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-mod…

Cited by 0SourceScholar