← Search

Shehzeen Hussain

4 accepted papers

2026

Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization

ICASSP 2026poster

Developing high-quality text-to-speech (TTS) systems for low-resource languages is challenging due to the scarcity of paired text and speech data. In contrast, automatic speech recognition (ASR) models for such languages are often more accessible, owing to large-scale multilingual pre-training effor…

Cited by 0SourcePDFScholar
2025

Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference

ICASSP 2025accepted

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and infe…

Cited by 0SourceScholar
2023

ACE-VC: Adaptive and Controllable Voice Conversion Using Explicitly Disentangled Self-Supervised Speech Representations

ICASSP 2023accepted

In this work, we propose a zero-shot voice conversion method using speech representations trained with self-supervised learning. First, we develop a multi-task model to decompose a speech utterance into features such as linguistic content, speaker characteristics, and speaking style. To disentangle…

Cited by 0SourceScholar
2022

Multi-Task Voice Activated Framework Using Self-Supervised Learning

ICASSP 2022accepted

Self-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition. Since these representations are learned without any task-specific supervision, they can also be u…

Cited by 0SourceScholar