← Search

Md Asif Jalal

3 accepted papers

2025

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

ICASSP 2025accepted

Diffusion based Text-To-Music (TTM) models generate music corresponding to text descriptions. Typically UNet based diffusion models condition on text embeddings generated from a pre-trained large language model or from a cross-modality audio-language representation model. This work proposes a diffus…

Cited by 0SourceScholar
2025

persoDA: Personalized Data Augmentation for Personalized ASR

ICASSP 2025accepted

Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word E…

Cited by 0SourceScholar
2023

Towards Domain Generalisation in ASR with Elitist Sampling and Ensemble Knowledge Distillation

ICASSP 2023accepted

Knowledge distillation (KD) has widely been used for model compression and domain adaptation for speech applications. In the presence of multiple teachers, knowledge can easily be transferred to the student by averaging the models output. However, previous research shows that the student do not adap…

Cited by 0SourceScholar