AISTATS 2017poster39 citations
Linear Convergence of Stochastic Frank Wolfe Variants
Donald Goldfarb, Garud Iyengar, Chaoxu Zhou
Abstract
In this paper, we show that the Away-step Stochastic Frank-Wolfe (ASFW) and Pairwise Stochastic Frank-Wolfe (PSFW) algorithms converge linearly in expectation. We also show that if an algorithm convergences linearly in expectation then it converges linearly almost surely. In order to prove these results, we develop a novel proof technique based on concepts of empirical processes and concentration inequalities. As far as we know, this technique has not been used previously to derive the convergence rates of stochastic optimization algorithms. In large- scale numerical experiments, ASFW and PSFW perform as well as or better than their stochastic competitors in actual CPU time.
BibTeX
@InProceedings{pmlr-v54-goldfarb17a,
title = {{Linear Convergence of Stochastic Frank Wolfe Variants}},
author = {Goldfarb, Donald and Iyengar, Garud and Zhou, Chaoxu},
booktitle = {Proceedings of the 20th International Conference on Artificial Intelligence and Statistics},
pages = {1066--1074},
year = {2017},
editor = {Singh, Aarti and Zhu, Jerry},
volume = {54},
series = {Proceedings of Machine Learning Research},
month = {20--22 Apr},
publisher = {PMLR},
pdf = {http://proceedings.mlr.press/v54/goldfarb17a/goldfarb17a.pdf},
url = {https://proceedings.mlr.press/v54/goldfarb17a.html},
abstract = {In this paper, we show that the Away-step Stochastic Frank-Wolfe (ASFW) and Pairwise Stochastic Frank-Wolfe (PSFW) algorithms converge linearly in expectation. We also show that if an algorithm convergences linearly in expectation then it converges linearly almost surely. In order to prove these results, we develop a novel proof technique based on concepts of empirical processes and concentration inequalities. As far as we know, this technique has not been used previously to derive the convergence rates of stochastic optimization algorithms. In large- scale numerical experiments, ASFW and PSFW perform as well as or better than their stochastic competitors in actual CPU time.}
}