← Search

Hemant A. Patil

12 accepted papers

2022

Constant Q Cepstral coefficients for classification of normal vs. Pathological infant cry

ICASSP 2022accepted

Classification of normal vs. pathological infant cry is an interesting and technologically challenging research problem due to quasi-periodic sampling of vocal tract spectrum by high pitch-source harmonics resulting in extremely poor spectral resolution for commonly used spectral features, such as M…

Cited by 0SourceScholar
2021

Cross-Teager Energy Cepstral Coefficients for Replay Spoof Detection on Voice Assistants

ICASSP 2021accepted

Voice assistants (VAs) are highly vulnerable to replay attacks, where the impostor plays pre-recorded voice samples to gain an unauthorized access to personalised devices. To that effect, we present an optimal microphone-channel selection scheme using Cross-Teager Energy Operator (CTEO) for spoofed…

Cited by 0SourceScholar
2020

Mspec-Net : Multi-Domain Speech Conversion Network

ICASSP 2020accepted

In this paper, we present a multi-domain speech conversion technique by proposing a Multi-domain Speech Conversion Network (MSpeC-Net) architecture for solving the less-explored area of Non-Audible Murmur-to-SPeeCH (NAM2-SPCH) conversion. The murmur produced by the speaker and captured by the NAM mi…

Cited by 0SourceScholar
2019

Analysis of Reverberation via Teager Energy Features for Replay Spoof Speech Detection

ICASSP 2019accepted

The Automatic Speaker Verification (ASV) systems are vulnerable to spoofing attacks. Detecting replay attack is the challenging Spoof Speech Detection (SSD) task, as several factors are involved during replay mechanism. Hence, it is important to analyze these factors for effective SSD task. This pap…

Cited by 0SourceScholar
2018

Time-Frequency Masking-Based Speech Enhancement Using Generative Adversarial Network

ICASSP 2018accepted

The success of time-frequency (T-F) mask-based approaches is dependent on the accuracy of predicted mask given the noisy spectral features. The state-of-the-art methods in T- F masking-based enhancement employ Deep Neural Network (DNN) to predict mask. Recently, Generative Adversarial Networks (GAN)…

Cited by 0SourceScholar
2017

Novel Amplitude Scaling method for bilinear frequency Warping-based Voice Conversion

ICASSP 2017accepted

In Frequency Warping (FW)-based Voice Conversion (VC), the source spectrum is modified to match the frequency-axis of the target spectrum followed by an Amplitude Scaling (AS) to compensate the amplitude differences between the warped spectrum and the actual target spectrum. In this paper, we propos…

Cited by 0SourceScholar
2017

Quality assessment of voice converted speech using articulatory features

ICASSP 2017accepted

We propose a novel application of the acoustic-to-articulatory inversion (AAI) towards a quality assessment of the voice converted speech. The ability of humans to speak effortlessly requires the coordinated movements of various articulators, muscles, etc. This effortless movement contributes toward…

Cited by 11SourceScholar
2016

Effectiveness of fundamental frequency (F0) and strength of excitation (SOE) for spoofed speech detection

ICASSP 2016accepted

Current countermeasures used in spoof detectors (for speech synthesis (SS) and voice conversion (VC)) are generally phase-based (as vocoders in SS and VC systems lack phase-information). These approaches may possibly fail for non-vocoder or unit-selection-based spoofs. In this work, we explore excit…

Cited by 0SourceScholar
2016

Filterbank learning using Convolutional Restricted Boltzmann Machine for speech recognition

ICASSP 2016accepted

Convolutional Restricted Boltzmann Machine (ConvRBM) as a model for speech signal is presented in this paper. We have developed ConvRBM with sampling from noisy rectified linear units (NReLUs). ConvRBM is trained in an unsupervised way to model speech signal of arbitrary lengths. Weights of the mode…

Cited by 0SourceScholar