2018
TETRIS: TilE-matching the TRemendous Irregular Sparsity
NeurIPS 2018poster
Compressing neural networks by pruning weights with small magnitudes can significantly reduce the computation and storage cost. Although pruning makes the model smaller, it is difficult to get practical speedup in modern computing platforms such as CPU and GPU due to the irregularity. Structural pru…