← Search

Hanbin Lee

3 accepted papers

2026

GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection

AAAI 2026technical

Knowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full a

Cited by 0SourcePDFScholar
2024

Learning Contextualized Representation on Discrete Space Via Hierarchical Product Quantization

ICASSP 2024accepted

Self-supervised learning has recently demonstrated significant success in various speech processing applications. Recent studies report that pre-training with contextualized continuous targets plays a crucial role in fine-tuning for better speech downstream tasks. However, unlike the continuous targ…

Cited by 0SourceScholar
2023

Masked Token Similarity Transfer for Compressing Transformer-Based ASR Models

ICASSP 2023accepted

Recent self-supervised automatic speech recognition (ASR) models based on transformers are showing best performance, but their footprint is too large to be trained on low-resource environments or deployed to edge devices. Knowledge distillation (KD) can be employed to reduce the model size. However,…

Cited by 0SourceScholar