← Search

Jaeho Lee*

2 accepted papers

2024

The Role of Masking for Efficient Supervised Knowledge Distillation of Vision Transformers

ECCV 2024poster

"Knowledge distillation is an effective method for training lightweight vision models. However, acquiring teacher supervision for training samples is often costly, especially from large-scale models like vision transformers (ViTs). In this paper, we develop a simple framework to reduce the supervisi…

Cited by 1SourcePDFScholar
2020

Lookahead: A Far-sighted Alternative of Magnitude-based Pruning

ICLR 2020poster

Magnitude-based pruning is one of the simplest methods for pruning neural networks. Despite its simplicity, magnitude-based pruning and its variants demonstrated remarkable performances for pruning modern architectures. Based on the observation that magnitude-based pruning indeed minimizes the Frobe…

Cited by 127SourcecodeScholar