← Search

Hyung Yong Kim

5 accepted papers

2026

GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection

AAAI 2026technical

Knowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full a

Cited by 0SourcePDFScholar
2024

Learning Contextualized Representation on Discrete Space Via Hierarchical Product Quantization

ICASSP 2024accepted

Self-supervised learning has recently demonstrated significant success in various speech processing applications. Recent studies report that pre-training with contextualized continuous targets plays a crucial role in fine-tuning for better speech downstream tasks. However, unlike the continuous targ…

Cited by 0SourceScholar
2023

Joint Unsupervised and Supervised Learning for Context-Aware Language Identification

ICASSP 2023accepted

Language identification (LID) recognizes the language of a spoken utterance automatically. According to recent studies, LID models trained with an automatic speech recognition (ASR) task perform better than those trained with a LID task only. However, we need additional text labels to train the mode…

Cited by 0SourceScholar
2023

Masked Token Similarity Transfer for Compressing Transformer-Based ASR Models

ICASSP 2023accepted

Recent self-supervised automatic speech recognition (ASR) models based on transformers are showing best performance, but their footprint is too large to be trained on low-resource environments or deployed to edge devices. Knowledge distillation (KD) can be employed to reduce the model size. However,…

Cited by 0SourceScholar
2020

Robust Front-End for Multi-Channel ASR using Flow-Based Density Estimation

IJCAI 2020poster

For multi-channel speech recognition, speech enhancement techniques such as denoising or dereverberation are conventionally applied as a front-end processor. Deep learning-based front-ends using such techniques require aligned clean and noisy speech pairs which are generally obtained via data simula…

Cited by 0SourcePDFScholar