← Search

Shirish S. Karande

3 accepted papers

2025

Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset

ICASSP 2025accepted

Current Non-Audible Murmur (NAM)-to-speech techniques rely on voice cloning to simulate ground-truth speech from paired whispers. However, the simulated speech often lacks intelligibility and fails to generalize well across different speakers. To address this issue, we focus on learning phoneme-leve…

Cited by 0SourceScholar
2025

HamaraAwaz: Advancing Low-Latency Streaming TTS for Multilingual Speech in Indian Languages

ICASSP 2025accepted

We present a multilingual, multi-speaker, low-latency speech synthesis system developed by the HamaraAwaz team for Track 1 of the LIMMITS’25 challenge. To improve speaker similarity and naturalness in Indic languages, we build on ParrotTTS. We utilize disentangled self-supervised speech representati…

Cited by 0SourceScholar
2025

MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI

ICASSP 2025accepted

Previous real-time MRI (rtMRI)-based speech synthesis models depend heavily on noisy ground-truth speech. Applying loss directly over ground truth mel-spectrograms entangles speech content with MRI noise, resulting in poor intelligibility. We introduce a novel approach that adapts the multi-modal se…

Cited by 0SourceScholar