ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive Skewness
Knowledge distillation (KD) is a widely adopted technique for transferring the capabilities of large teacher models to smaller student models, thereby significantly reducing inference costs and memory consumption. However, existing KD methods are all constrained by an inherent greedy optimization ob