2025
Effective Interplay between Sparsity and Quantization: From Theory to Practice
ICLR 2025spotlight
The increasing size of deep neural networks (DNNs) necessitates effective model compression to reduce their computational and memory footprints. Sparsity and quantization are two prominent compression methods that have been shown to reduce DNNs' computational and memory footprints significantly whil…