← Search

Srinivasan Umesh

10 accepted papers

2025

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

EMNLP 2025

Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectures still tend to struggle in zero-shot cross-lingual settings. They are also often unable to generalize for speakers of

2024

FusDom: Combining in-Domain and Out-of-Domain Knowledge for Continuous Self-Supervised Learning

ICASSP 2024accepted

Continued pre-training (CP) offers multiple advantages, like target domain adaptation and the potential to exploit the continuous stream of unlabeled data available online. However, continued pre-training on out-of-domain distributions often leads to catastrophic forgetting of previously acquired kn…

Cited by 0SourceScholar
2024

Stable Distillation: Regularizing Continued Pre-Training for Low-Resource Automatic Speech Recognition

ICASSP 2024accepted

Continued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain has shown to be extremely effective for low-resource Automatic Speech Recognition (ASR). This paper proposes Stable Distillation, a simple and novel approach for SSL-based continued pre-training that b…

Cited by 0SourceScholar
2023

Data2vec-Aqc: Search for the Right Teaching Assistant in the Teacher-Student Training Setup

ICASSP 2023accepted

In this paper, we propose a new Self-Supervised Learning (SSL) algorithm called data2vec-aqc, for speech representation learning from unlabeled speech data. Our goal is to improve SSL for speech in domains where both unlabeled and labeled data are limited. Building on the recently introduced data2ve…

Cited by 0SourceScholar
2023

SLICER: Learning Universal Audio Representations Using Low-Resource Self-Supervised Pre-Training

ICASSP 2023accepted

We present a new Self-Supervised Learning (SSL) approach to pre-train encoders on unlabeled audio data that reduces the need for large amounts of labeled data for audio and speech classification. Our primary aim is to learn au-dio representations that can generalize across a large vari-ety of speech…

Cited by 0SourceScholar
2022

Investigation of Robustness of Hubert Features from Different Layers to Domain, Accent and Language Variations

ICASSP 2022accepted

In this paper, we investigate the use of pre-trained HuBERT model to build downstream Automatic Speech Recognition (ASR) models using data that have differences in domain, accent and even language. We use the standard ESPnet recipe with HuBERT as pre-trained models whose output is fed as input featu…

Cited by 0SourceScholar
2021

Exploring the use of Common Label Set to Improve Speech Recognition of Low Resource Indian Languages

ICASSP 2021accepted

In many Indian languages, written characters are organized on sound phonetic principles, and the ordering of characters is the same across many of them. However, while training conventional end-to-end (E2E) Multilingual speech recognition systems, we treat characters or target subword units from dif…

Cited by 16SourceScholar
2020

Improving the Performance of Transformer Based Low Resource Speech Recognition for Indian Languages

ICASSP 2020accepted

The recent success of the Transformer based sequence-to-sequence framework for various Natural Language Processing tasks has motivated its application to Automatic Speech Recognition. In this work, we explore the application of Transformers on low resource Indian languages in a multilingual framewor…

Cited by 0SourceScholar
2020

Investigation of Methods to Improve the Recognition Performance of Tamil-English Code-Switched Data in Transformer Framework

ICASSP 2020accepted

Code-switching (CS) refers to (inter/intra-word) switching between multiple languages in a single conversation. In multilingual countries like India, CS occurs very often in everyday speech, resulting in a new breed of languages in urban regions like Hinglish (Hindi-English), Tanglish (Tamil-English…

Cited by 22SourceScholar