← Search

Trambak Banerjee

1 accepted papers

2021

Rethinking gradient sparsification as total error minimization

NeurIPS 2021spotlight

Gradient compression is a widely-established remedy to tackle the communication bottleneck in distributed training of large deep neural networks (DNNs). Under the error-feedback framework, Top-$k$ sparsification, sometimes with $k$ as little as 0.1% of the gradient size, enables training to the same…

Cited by 68SourcePDFScholar