← Search

Vimal Manohar

10 accepted papers

2025

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

ICLR 2025poster

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre and style, leading to challenges in achieving controllable g…

2024

Less Peaky and More Accurate CTC Forced Alignment by Label Priors

ICASSP 2024accepted

Connectionist temporal classification (CTC) models are known to have peaky output distributions. Such behavior is not a problem for automatic speech recognition (ASR), but it can cause inaccurate forced alignments (FA), especially at finer granularity, e.g., phoneme level. This paper aims at allevia…

Cited by 0SourceScholar
2023

Self-Supervised Representations for Singing Voice Conversion

ICASSP 2023accepted

A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Vec 2.0 have helped further the state-of-the-art. Though these methods produce mor…

Cited by 25SourceScholar
2023

Voice-Preserving Zero-Shot Multiple Accent Conversion

ICASSP 2023accepted

Most people who have tried to learn a foreign language would have experienced difficulties understanding or speaking with a native speaker’s accent. For native speakers, understanding or speaking a new accent is likewise a difficult task. An accent conversion system that changes a speaker’s accent b…

Cited by 26SourceScholar
2023

Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

NeurIPS 2023poster

Large-scale generative models such as GPT and DALL-E have revolutionized the research community. These models not only generate high fidelity outputs, but are also generalists which can solve tasks not explicitly taught. In contrast, speech generative models are still primitive in terms of scale and…

Cited by 299SourcePDFScholar
2019

Acoustic Modeling for Overlapping Speech Recognition: Jhu Chime-5 Challenge System

ICASSP 2019accepted

This paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party speech recorded by multiple microphone arrays. We explore data augmentation approaches, neural network architectures, fr…

Cited by 0SourceScholar
2019

Towards Automatic Methods to Detect Errors in Transcriptions of Speech Recordings

ICASSP 2019accepted

This work explores different methods to detect errors in transcriptions of speech recordings. We artificially corrupt well transcribed speech transcriptions with three types of errors: substitution, insertion and deletion on TIMIT phonemic transcriptions and WSJ word transcriptions. First, we use Ba…

Cited by 0SourceScholar
2018

Characterizing Performance of Speaker Diarization Systems on Far-Field Speech Using Standard Methods

ICASSP 2018accepted

To date, the bulk of research on speaker diarization has been conducted on telephone or near-field speech. As the need for technologies capable of handling conversational speech increases, it is necessary to establish the performance of state-of-the-art systems in this domain. In this work we evalua…

Cited by 0SourceScholar
2018

Semi-Supervised Training of Acoustic Models Using Lattice-Free MMI

ICASSP 2018accepted

The lattice-free MMI objective (LF-MMI) has been used in supervised training of state-of-the-art neural network acoustic models for automatic speech recognition (ASR). With large amounts of unsupervised data available, extending this approach to the semi-supervised scenario is of significance. Finit…

Cited by 0SourceScholar
2016

Adapting ASR for under-resourced languages using mismatched transcriptions

ICASSP 2016accepted

Mismatched transcriptions of speech in a target language refers to transcriptions provided by people unfamiliar with the language, using English letter sequences. In this work, we demonstrate the value of such transcriptions in building an ASR system for the target language. For different languages,…

Cited by 0SourceScholar