2025
Adaptive Pruning of Pretrained Transformer via Differential Inclusions
ICLR 2025poster
Large transformers have demonstrated remarkable success, making it necessary to compress these models to reduce inference costs while preserving their performance. Current compression algorithms prune transformers at fixed compression ratios, requiring a unique pruning process for each ratio, which…