← Search

Suyoun Kim

7 accepted papers

2024

PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding

EMNLP 2024finding

Spoken Language Understanding (SLU) is a critical component of voice assistants; it consists of converting speech to semantic parses for task execution. Previous works have explored end-to-end models to improve the quality and robustness of SLU models with Deliberation, however these models have rem…

Cited by 0SourcePDFScholar
2023

ICASSP 2023 Spoken Language Understanding Grand Challenge

ICASSP 2023accepted

Spoken language understanding (SLU) is a important field between the Speech and NLP community focused on converting a users’ speech utterance into an executable semantic parse. In order to facilitate open research in this space, we introduce the 1st Spoken Language Understanding challenge hosted at…

Cited by 0SourceScholar
2023

Introducing Semantics into Speech Encoders

ACL 2023long

Recent studies find existing self-supervised speech encoders contain primarily acoustic rather than semantic information. As a result, pipelined supervised automatic speech recognition (ASR) to large language model (LLM) systems achieve state-of-the-art results on semantic spoken language tasks by u…

Cited by 4SourcePDFScholar
2022

Joint Audio/Text Training for Transformer Rescorer of Streaming Speech Recognition

EMNLP 2022finding

Recently, there has been an increasing interest in two-pass streaming end-to-end speech recognition (ASR) that incorporates a 2nd-pass rescoring model on top of the conventional 1st-pass streaming ASR model to improve recognition accuracy while keeping latency low. One of the latest 2nd-pass rescori…

Cited by 8SourcePDFScholar
2021

Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer

ICASSP 2021accepted

Recurrent Neural Network Transducer (RNN-T), like most end-to-end speech recognition model architectures, has an implicit neural network language model (NNLM) and cannot easily leverage unpaired text data during training. Previous work has proposed various fusion methods to incorporate external NNLM…

Cited by 0SourceScholar
2017

Joint CTC-attention based end-to-end speech recognition using multi-task learning

ICASSP 2017accepted

Recently, there has been an increasing interest in end-to-end speech recognition that directly transcribes speech to text without any predefined alignments. One approach is the attention-based encoder-decoder framework that learns a mapping between variable-length input and output sequences in one s…

Cited by 0SourceScholar