← Search

Alexander Olshevsky

4 accepted papers

2021

Communication-efficient SGD: From Local SGD to One-Shot Averaging

NeurIPS 2021poster

We consider speeding up stochastic gradient descent (SGD) by parallelizing it across multiple workers. We assume the same data set is shared among $N$ workers, who can take SGD steps and coordinate with a central server. While it is possible to obtain a linear reduction in the variance by averaging…

Cited by 29SourcePDFScholar
2019

Graph Resistance and Learning from Pairwise Comparisons

ICML 2019oral

We consider the problem of learning the qualities of a collection of items by performing noisy comparisons among them. Following the standard paradigm, we assume there is a fixed “comparison graph” and every neighboring pair of items in this graph is compared k times according to the Bradley-Terry-L…

Cited by 15SourcePDFScholar
2018

Gradient Descent for Sparse Rank-One Matrix Completion for Crowd-Sourced Aggregation of Sparsely Interacting Workers

ICML 2018oral

We consider worker skill estimation for the single coin Dawid-Skene crowdsourcing model. In practice skill-estimation is challenging because worker assignments are sparse and irregular due to the arbitrary, and uncontrolled availability of workers. We formulate skill estimation as a rank-one correla…

Cited by 31SourcePDFScholar