2026
The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMM
ICLR 2026poster
Neural network pruning is a promising technique to mitigate the excessive computational and memory requirements of large language models (LLMs). Despite its promise, however, progress in this area has diminished, as conventional methods are seemingly unable to surpass moderate sparsity levels (50-60…