← Search

Gary Wang

6 accepted papers

2025

Speech Re-Painting for Robust ASR

ICASSP 2025accepted

Synthetic speech is a useful source for augmentation of automatic speech recognition (ASR) systems, but there is a "sim-to-real" gap between synthetic and real speech that can limit generalization. The natural variability of real speech is essential to the training of robust ASR systems. While synth…

Cited by 0SourceScholar
2024

Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data

ICASSP 2024accepted

Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data without supervision. The proposed framework combines speech-t…

Cited by 0SourceScholar
2023

Understanding Shared Speech-Text Representations

ICASSP 2023accepted

Recently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting sh…

Cited by 0SourceScholar
2023

Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-to-Speech

ICASSP 2023accepted

This paper proposes Virtuoso, a massively multilingual speech–text joint semi-supervised learning framework for text-to-speech synthesis (TTS) models. Existing multilingual TTS typically supports tens of languages, which are a small fraction of the thousands of languages in the world. One difficulty…

Cited by 0SourceScholar
2022

Tts4pretrain 2.0: Advancing the use of Text and Speech in ASR Pretraining with Consistency and Contrastive Losses

ICASSP 2022accepted

An effective way to learn representations from untranscribed speech and unspoken text with linguistic/lexical representations derived from synthesized speech was introduced in tts4pretrain [1]. However, the representations learned from synthesized and real speech are likely to be different, potentia…

Cited by 0SourceScholar
2020

Improving Speech Recognition Using Consistent Predictions on Synthesized Speech

ICASSP 2020accepted

Speech synthesis has advanced to the point of being close to indistinguishable from human speech. However, efforts to train speech recognition systems on synthesized utterances have not been able to show that synthesized data can be effectively used to augment or replace human speech. In this work,…

Cited by 0SourceScholar