2023
Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
ICASSP 2023accepted
Knowledge distillation (KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour of a teacher model. However, traditional KD methods suffer from teacher label storage issue, especially when the train…