← Search

Mahesh Kumar Nandwana

5 accepted papers

2026

REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation

AAAI 2026technical

Simultaneous Speech Translation (SimulST) systems stream in audio while simultaneously emitting translated text or speech. Such systems face the significant challenge of balancing translation quality and latency. We introduce a strategy to optimize this tradeoff: wait for more input only if you gain

Cited by 0SourcePDFScholar
2024

Voice Toxicity Detection Using Multi-Task Learning

ICASSP 2024accepted

Social communication systems must identify toxic voice audio to support moderation that protects the safety and civility of their communities. Toxicity classification for voice depends on both audio style, such as volume and tone, and content, such as the words in the speech individually and in cont…

Cited by 0SourceScholar
2019

Analysis and Mitigation of Vocal Effort Variations in Speaker Recognition

ICASSP 2019accepted

In this work, we assess the impact of vocal effort on discrimination and calibration performance of a state-of-the-art speaker recognition system. We analyze three levels of vocal effort (low, normal, and high) from the SRI-FRTIV corpus. We use a deep neural network (DNN) speaker embeddings system w…

Cited by 0SourceScholar
2016

Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach

ICASSP 2016accepted

Sustaining automatic speaker verification(ASV) systems from spoofing attacks remains an essential challenge, even if significant progress in ASV has been achieved in recent years. In this study, an automatic spoofing detection approach using an i-vector framework is proposed. Two approaches are used…

Cited by 0SourceScholar
2015

Robust unsupervised detection of human screams in noisy acoustic environments

ICASSP 2015accepted

This study is focused on an unsupervised approach for detection of human scream vocalizations from continuous recordings in noisy acoustic environments. The proposed detection solution is based on compound segmentation, which employs weighted mean distance, T <sup xmlns:mml="http://www.w3.org/1998/M…

Cited by 0SourceScholar