← Search

Shukjae Choi

8 accepted papers

2026

GrayKD: Distilling Better Knowledge from Black-box LLM via Multi-rationale Injection

AAAI 2026technical

Knowledge distillation (KD) is a promising compression technique for reducing the computational burden of large language models (LLMs). Depending on access to the teacher model’s internal parameters, KD is typically categorized into white-box and black-box KD. While white-box KD benefits from full a

Cited by 0SourcePDFScholar
2026

SPADE: STRUCTURED PRUNING AND ADAPTIVE DISTILLATION FOR EFFICIENT LLM-TTS

ICASSP 2026oral

The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent LLM-TTS systems achieve strong controllability and zero-shot generalization, but their large parameter counts and high…

Cited by 0SourcePDFScholar
2025

Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding

ICASSP 2025accepted

The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to pre…

Cited by 0SourceScholar
2024

Boosting Unknown-Number Speaker Separation with Transformer Decoder-Based Attractor

ICASSP 2024accepted

We propose a novel speech separation model designed to separate mixtures with an unknown number of speakers. The proposed model stacks 1) a dual-path processing block that can model spectro-temporal patterns, 2) a transformer decoder-based attractor (TDA) calculation module that can deal with an unk…

Cited by 0SourceScholar
2024

VoxtLM: Unified Decoder-Only Models for Consolidating Speech Recognition, Synthesis and Speech, Text Continuation Tasks

ICASSP 2024accepted

We propose a decoder-only language model, VoxtLM, that can perform four tasks: speech recognition, speech synthesis, text generation, and speech continuation. VoxtLM integrates text vocabulary with discrete speech tokens from self-supervised speech features and uses special tokens to enable multitas…

Cited by 99SourceScholar
2023

FNeural Speech Enhancement with Very Low Algorithmic Latency and Complexity via Integrated full- and sub-band Modeling

ICASSP 2023accepted

We propose FSB-LSTM, a novel long short-term memory (LSTM) based architecture that integrates full- and sub-band (FSB) modeling, for single- and multi-channel speech enhancement in the short-time Fourier transform (STFT) domain. The model maintains an information highway to flow an over-complete inp…

Cited by 0SourceScholar
2023

Joint Unsupervised and Supervised Learning for Context-Aware Language Identification

ICASSP 2023accepted

Language identification (LID) recognizes the language of a spoken utterance automatically. According to recent studies, LID models trained with an automatic speech recognition (ASR) task perform better than those trained with a LID task only. However, we need additional text labels to train the mode…

Cited by 0SourceScholar
2023

TF-GRIDNET: Making Time-Frequency Domain Models Great Again for Monaural Speaker Separation

ICASSP 2023accepted

We propose TF-GridNet, a novel multi-path deep neural network (DNN) operating in the time-frequency (T-F) domain, for monaural talker-independent speaker separation in anechoic conditions. The model stacks several multi-path blocks, each consisting of an intra-frame spectral module, a sub-band tempo…

Cited by 182SourceScholar