← Search

Byeong-Yeol Kim

9 accepted papers

2026

GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection

AAAI 2026technical

Knowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full a

Cited by 0SourcePDFScholar
2024

Boosting Unknown-Number Speaker Separation with Transformer Decoder-Based Attractor

ICASSP 2024accepted

We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can model spectro-temporal patterns, 2) a transformer decoder-based attractor (TDA) calculation module that can deal with an unk…

Cited by 0SourceScholar
2024

Faces that Speak: Jointly Synthesising Talking Face and Speech from Text

CVPR 2024poster

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the main challenges of each task: (1) generating a range of head…

Cited by 11SourcePDFScholar
2024

Learning Contextualized Representation on Discrete Space Via Hierarchical Product Quantization

ICASSP 2024accepted

Self-supervised learning has recently demonstrated significant success in various speech processing applications. Recent studies report that pre-training with contextualized continuous targets plays a crucial role in fine-tuning for better speech downstream tasks. However, unlike the continuous targ…

Cited by 0SourceScholar
2023

FNeural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated full- and sub-band Modeling

ICASSP 2023accepted

We propose FSB-LSTM, a novel long short-term memory (LSTM) based architecture that integrates full- and sub-band (FSB) modeling, for single- and multi-channel speech enhancement in the short-time Fourier transform (STFT) domain. The model maintains an information highway to flow an over-complete inp…

Cited by 0SourceScholar
2023

Joint Unsupervised and Supervised Learning for Context-Aware Language Identification

ICASSP 2023accepted

Language identification (LID) recognizes the language of a spoken utterance automatically. According to recent studies, LID models trained with an automatic speech recognition (ASR) task perform better than those trained with a LID task only. However, we need additional text labels to train the mode…

Cited by 0SourceScholar
2023

Masked Token Similarity Transfer for Compressing Transformer-Based ASR Models

ICASSP 2023accepted

Recent self-supervised automatic speech recognition (ASR) models based on transformers are showing best performance, but their footprint is too large to be trained on low-resource environments or deployed to edge devices. Knowledge distillation (KD) can be employed to reduce the model size. However,…

Cited by 0SourceScholar
2023

Metric Learning for User-Defined Keyword Spotting

ICASSP 2023accepted

The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits their transferability to unseen terms. The ability to define custom keywords has advantages in terms of user experience.I…

Cited by 0SourceScholar
2023

TF-GRIDNET: Making Time-Frequency Domain Models Great Again for Monaural Speaker Separation

ICASSP 2023accepted

We propose TF-GridNet, a novel multi-path deep neural network (DNN) operating in the time-frequency (T-F) domain, for monaural talker-independent speaker separation in anechoic conditions. The model stacks several multi-path blocks, each consisting of an intra-frame spectral module, a sub-band tempo…

Cited by 182SourceScholar