← Search

Kyu Jeong Han

5 accepted papers

2023

Wav2Seq: Pre-Training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages

ICASSP 2023accepted

We introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervised pseudo speech recognition task — transcribing audio inputs into pseudo subword…

Cited by 0SourceScholar
2022

Performance-Efficiency Trade-Offs in Unsupervised Pre-Training for Speech Recognition

ICASSP 2022accepted

This paper is a study of performance-efficiency trade-offs in pre-trained models for automatic speech recognition (ASR). We focus on wav2vec 2.0, and formalize several architecture designs that influence both the model performance and its efficiency. Putting together all our observations, we introdu…

Cited by 0SourceScholar
2022

SLUE: New Benchmark Tasks For Spoken Language Understanding Evaluation on Natural Speech

ICASSP 2022accepted

Progress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in higher-level spoken language understanding tasks, including us…

Cited by 0SourceScholar
2022

SRU++: Pioneering Fast Recurrence with Attention for Speech Recognition

ICASSP 2022accepted

The Transformer architecture has been well adopted as a dominant architecture in most sequence transduction tasks including automatic speech recognition (ASR), since its attention mechanism excels in capturing long-range dependencies. While models built solely upon attention can be better paralleliz…

Cited by 0SourceScholar
2021

Multistream CNN for Robust Acoustic Modeling

ICASSP 2021accepted

This paper proposes multistream CNN, a novel neural network architecture for robust acoustic modeling in speech recognition tasks. The proposed architecture processes input speech with diverse temporal resolutions by applying different dilation rates to convolutional neural networks across multiple…

Cited by 0SourceScholar