← Search

Hong-Kwang Jeff Kuo

9 accepted papers

2025

LLM based Text Generation for Improved Low-resource Speech Recognition Models

ICASSP 2025accepted

Limited transcribed spoken style data is a critical bottleneck in building automatic speech recognition (ASR) systems for low-resource languages. Prompting a large language model (LLM) to paraphrase input text can generate novel text data that is constrained to be semantically similar to the source…

Cited by 0SourceScholar
2023

Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and Understanding

ICASSP 2023accepted

RNN Tranducer (RNN-T) technology is very popular for building deployable models for end-to-end (E2E) automatic speech recognition (ASR) and spoken language understanding (SLU). Since these are E2E models operating on speech directly, there remains a potential to improve their performance using purel…

Cited by 0SourceScholar
2023

Multi-Speaker Data Augmentation for Improved end-to-end Automatic Speech Recognition

ICASSP 2023accepted

Publicly available datasets traditionally used to train E2E ASR models for conversational telephone speech recognition are based on clean, short duration, single speaker utterances collected on separate channels. While E2E ASR models achieve state-of-the-art performance on recognition tasks that mat…

Cited by 0SourceScholar
2022

Improving End-to-end Models for Set Prediction in Spoken Language Understanding

ICASSP 2022accepted

The goal of spoken language understanding (SLU) systems is to determine the meaning of the input speech signal, unlike speech recognition which aims to produce verbatim transcripts. Advances in end-to-end (E2E) speech modeling have made it possible to train solely on semantic entities, which are far…

Cited by 0SourceScholar
2022

Integrating Text Inputs for Training and Adapting RNN Transducer ASR Models

ICASSP 2022accepted

Compared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be in-dependently adapted to a new domain, recent end-to-end (E2E) ASR system are harder to customize due to their all-neural monolithic construction. In this paper, we propose a…

Cited by 0SourceScholar
2022

Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding

ICASSP 2022accepted

Dialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in text form, which makes the model dependent on a cascaded automatic speech recognizer (ASR). This rescinds the benefits of a…

Cited by 0SourceScholar
2022

Towards Reducing the Need for Speech Training Data to Build Spoken Language Understanding Systems

ICASSP 2022accepted

The lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly process speech inputs. In contrast, large amounts of text data with suitable labels are usually available. In this paper, we p…

Cited by 0SourceScholar
2021

End-to-End Spoken Language Understanding Using Transformer Networks and Self-Supervised Pre-Trained Features

ICASSP 2021accepted

Transformer networks and self-supervised pre-training have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of spoken language understanding (SLU) still need further investigation. In this paper we introduce a modular E…

Cited by 0SourceScholar
2021

RNN Transducer Models for Spoken Language Understanding

ICASSP 2021accepted

We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available ann…

Cited by 0SourceScholar