Theoretical Guarantees for One-Shot Magnitude Pruning and Compute-Adaptive Early Exit
We study compute reduction in neural networks through a unified partial versus full computation view, captured by one-shot magnitude pruning in the static regime and early exit in the adaptive regime. In an asymptotic single-neuron model, we prove a concentration theorem for one-shot magnitude pruni…