← Search

Yu-An Chung

11 accepted papers

2024

COLLD: Contrastive Layer-to-Layer Distillation for Compressing Multilingual Pre-Trained Speech Encoders

ICASSP 2024accepted

Large-scale self-supervised pre-trained speech encoders outperform conventional approaches in speech recognition and translation tasks. Due to the high cost of developing these large models, building new encoders for new tasks and deploying them to on-device applications are infeasible. Prior studie…

Cited by 0SourceScholar
2023

Speech-to-Speech Translation for a Real-world Unwritten Language

ACL 2023findings

We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard text writing systems. We use English-Taiwanese Hokkien as a case study, and present an end-to-end solution from training d…

2023

UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units

ACL 2023long

Direct speech-to-speech translation (S2ST), in which all components can be optimized jointly, is advantageous over cascaded approaches to achieve fast inference with a simplified pipeline. We present a novel two-pass direct S2ST architecture, UnitY, which first generates textual representations and…

2022

SSAST: Self-Supervised Audio Spectrogram Transformer

AAAI 2022technical

Recently, neural networks based purely on self-attention, such as the Vision Transformer (ViT), have been shown to outperform deep learning models constructed with convolutional neural networks (CNNs) on various vision tasks, thus extending the success of Transformers, which were originally develope…

2021

SPLAT: Speech-Language Joint Pre-Training for Spoken Language Understanding

NAACL 2021long

Spoken language understanding (SLU) requires a model to analyze input acoustic signal to understand its linguistic content and make predictions. To boost the models’ performance, various pre-training methods have been proposed to learn rich representations from large-scale unannotated speech and tex…

Cited by 81SourcePDFScholar
2019

Disentangling Correlated Speaker and Noise for Speech Synthesis via Data Augmentation and Adversarial Factorization

ICASSP 2019accepted

To leverage crowd-sourced data to train multi-speaker text-to-speech (TTS) models that can synthesize clean speech for all speakers, it is essential to learn disentangled representations which can independently control the speaker identity and background noise in generated signals. However, learning…

Cited by 0SourceScholar
2019

Semi-supervised Training for Improving Data Efficiency in End-to-end Speech Synthesis

ICASSP 2019accepted

Although end-to-end text-to-speech (TTS) models such as Tacotron have shown excellent results, they typically require a sizable set of high-quality <;text, audio> pairs for training, which are expensive to collect. In this paper, we propose a semi-supervised training framework to improve the data ef…

Cited by 0SourceScholar
2018

Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces

NeurIPS 2018spotlight

Recent research has shown that word embedding spaces learned from text corpora of different languages can be aligned without any parallel data supervision. Inspired by the success in unsupervised cross-lingual word embeddings, in this paper we target learning a cross-modal alignment between the embe…

Cited by 116SourcePDFScholar