2021
Reinforced Multi-Teacher Selection for Knowledge Distillation
AAAI 2021technical
In natural language processing (NLP) tasks, slow inference speed and huge footprints in GPU usage remain the bottleneck of applying pre-trained deep models in production. As a popular method for model compression, knowledge distillation transfers knowledge from one or multiple large (teacher) models…