← Search

Naomi Sagan

1 accepted papers

2024

Compressing Large Language Models using Low Rank and Low Precision Decomposition

NeurIPS 2024poster

The prohibitive sizes of Large Language Models (LLMs) today make it difficult to deploy them on memory-constrained edge devices. This work introduces $\rm CALDERA$ -- a new post-training LLM compression algorithm that harnesses the inherent low-rank structure of a weight matrix $\mathbf{W}$ by appro…