← Search

William Dally

2 accepted papers

2022

Optimal Clipping and Magnitude-aware Differentiation for Improved Quantization-aware Training

ICML 2022spotlight

Data clipping is crucial in reducing noise in quantization operations and improving the achievable accuracy of quantization-aware training (QAT). Current practices rely on heuristics to set clipping threshold scalars and cannot be shown to be optimal. We propose Optimally Clipped Tensors And Vectors…

Cited by 45SourcePDFScholar
2015

Learning both Weights and Connections for Efficient Neural Network

NeurIPS 2015poster

Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems. Also, conventional networks fix the architecture before training starts; as a result, training cannot improve the architecture. To address these limitations, we describe a me…

Cited by 8960SourcePDFScholar