2026
Beyond Student: An Asymmetric Network for Neural Network Inheritance
ICLR 2026poster
Knowledge Distillation (KD) has emerged as a powerful technique for model compression, enabling lightweight student networks to benefit from the performance of redundant teacher networks. However, the inherent capacity gap often limits the performance of student networks. Inspired by the expressiven…