2026
PADD: Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student Learning
ICML 2026poster
As large language models (LLMs) continue to scale, it becomes increasingly challenging to grow model capacity under fixed computation budgets. We propose Path-Aligned Decompression Distillation (PADD), a framework for distilling knowledge from dense teachers without explicit routing into mixture-of-…