← Search

Pablo Peso Parada

5 accepted papers

2025

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

ICASSP 2025accepted

Diffusion based Text-To-Music (TTM) models generate music corresponding to text descriptions. Typically UNet based diffusion models condition on text embeddings generated from a pre-trained large language model or from a cross-modality audio-language representation model. This work proposes a diffus…

Cited by 0SourceScholar
2025

Retrieval Augmented Generation based context discovery for ASR

EMNLP 2025

This work investigates retrieval augmented generation as an efficient strategy for automatic context discovery in context-aware Automatic Speech Recognition (ASR) system, in order to improve transcription accuracy in the presence of rare or out-of-vocabulary terms. However, identifying the right con

Cited by 0SourcePDFScholar
2025

ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization

ICASSP 2025accepted

Automatic Speech Recognition (ASR) is widely used within consumer devices such as mobile phones. Recently, personalization or on-device model fine-tuning has shown that adaptation of ASR models towards target user speech improves their performance over rare words or accented speech. Despite these ga…

Cited by 0SourceScholar
2025

persoDA: Personalized Data Augmentation for Personalized ASR

ICASSP 2025accepted

Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word E…

Cited by 0SourceScholar