2023
Progressive Ensemble Distillation: Building Ensembles for Efficient Inference
NeurIPS 2023poster
Knowledge distillation is commonly used to compress an ensemble of models into a single model. In this work we study the problem of progressive ensemble distillation: Given a large, pretrained teacher model , we seek to decompose the model into an ensemble of smaller, low-inference cost student mode…