2026
SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines
ICLR 2026poster
Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs. While KD has shown strong empirical success in numerous applications, its theoretical underpinnings remain o…