2021
Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation
IJCAI 2021poster
Knowledge distillation (KD), transferring knowledge from a cumbersome teacher model to a lightweight student model, has been investigated to design efficient neural architectures. Generally, the objective function of KD is the Kullback-Leibler (KL) divergence loss between the softened probability di…