2020
Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient Space
NeurIPS 2020poster
Distilling knowledge from an ensemble of teacher models is expected to have a more promising performance than that from a single one. Current methods mainly adopt a vanilla average rule, i.e., to simply take the average of all teacher losses for training the student network. However, this approach t…