← Search

Huiyuan Tian

2 accepted papers

2026

Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers

AAAI 2026technical

While feature-based knowledge distillation has proven highly effective for compressing CNNs, these techniques unexpectedly fail when applied to Vision Transformers (ViTs), often performing worse than simple logit-based distillation. We provide the first comprehensive analysis of this phenomenon thro

Cited by 0SourcePDFScholar
2026

From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers

ICML 2026poster

Feature-map knowledge distillation (KD) transfers internal representations well between comparably sized Vision Transformers (ViTs), but it often fails in compression. We revisit this failure and uncover a paradox. Sample-wise SVD shows that each image is highly compressible, which seems to suggest …

Cited by 0SourceScholar