← Search

Vinayak Tantia

2 accepted papers

2020

Lookahead Converges to Stationary Points of Smooth Non-convex Functions

ICASSP 2020accepted

The Lookahead optimizer [Zhang et al., 2019] was recently proposed and demonstrated to improve performance of stochastic first-order methods for training deep neural networks. Lookahead can be viewed as a two time-scale algorithm, where the fast dynamics (inner optimizer) determine a search directio…

Cited by 0SourceScholar
2020

SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum

ICLR 2020poster

Distributed optimization is essential for training large models on large datasets. Multiple approaches have been proposed to reduce the communication overhead in distributed training, such as synchronizing only after performing multiple local SGD steps, and decentralized methods (e.g., using gossip…

Cited by 217SourcecodeScholar