← Search

Kyu J. Han

6 accepted papers

2025

Contextual ASR with Retrieval Augmented Large Language Model

ICASSP 2025accepted

Automatic speech recognition (ASR) systems can benefit from incorporating contextual information to improve recognition accuracy, especially for uncommon words or phrases. Current approaches like custom vocabularies or prompting with previous transcript segments provide limited contextual control. C…

Cited by 0SourceScholar
2025

Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation

NAACL 2025long

Faithfulness evaluators based on Large Language Models (LLMs) are often fooled by the fluency of the text and struggle with identifying errors in the summaries, usually leading to high false negative rate. We propose an approach to summary faithfulness evaluation in which multiple LLM-based agents a…

2025

Knowledge Distillation From Ensemble for Spoken Language Identification

ICASSP 2025accepted

Spoken language identification (LID) has seen substantial performance gains with the rise of large-scale models. However, these models are often computationally expensive and impractical for many real-world applications. In this work, we propose a novel knowledge distillation from ensemble framework…

Cited by 0SourceScholar
2025

Speech Retrieval-Augmented Generation without Automatic Speech Recognition

ICASSP 2025accepted

One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. While this cascaded pipeline has proven effective in many practical settings, ASR…

Cited by 17SourceScholar
2025

Zero-resource Speech Translation and Recognition with LLMs

ICASSP 2025accepted

Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never…

Cited by 0SourceScholar