2024
Improving Knowledge Distillation via Regularizing Feature Direction and Norm
ECCV 2024oral
"Knowledge distillation (KD) is a particular technique of model compression that exploits a large well-trained teacher neural network to train a small student network . Treating teacher’s feature as knowledge, prevailing methods train student by aligning its features with the teacher’s, e.g., by min…