2025
SDPGO: Efficient Self-Distillation Training Meets Proximal Gradient Optimization
NeurIPS 2025poster
Self-knowledge distillation (SKD) enables single-model training by distilling knowledge from the model's own output, eliminating the need for a separate teacher network required in conventional distillation methods. However, current SKD methods focus mainly on replicating common features in the stud…