← Search

Yaya Sy

2 accepted papers

2025

Efficient One-shot Compression via Low-Rank Local Feature Distillation

NAACL 2025long

Current structured pruning approaches for large language models typically involve two steps: (1) compression using calibration data and (2) costly continued pretraining on billions of tokens to recover lost performance. This second step is necessary as the first significantly impacts model accuracy.…