ICML 2020poster111 citations

The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture

Francesca Mignacco, Florent Krzakala, Yue Lu, Pierfrancesco Urbani, Lenka Zdeborova

Abstract

We consider a high-dimensional mixture of two Gaussians in the noisy regime where even an oracle knowing the centers of the clusters misclassifies a small but finite fraction of the points. We provide a rigorous analysis of the generalization error of regularized convex classifiers, including ridge, hinge and logistic regression, in the high-dimensional limit where the number $n$ of samples and their dimension $d$ go to infinity while their ratio is fixed to $\alpha=n/d$. We discuss surprising effects of the regularization that in some cases allows to reach the Bayes-optimal performances. We also illustrate the interpolation peak at low regularization, and analyze the role of the respective sizes of the two clusters.

BibTeX
@InProceedings{pmlr-v119-mignacco20a,
  title = 	 {The Role of Regularization in Classification of High-dimensional Noisy {G}aussian Mixture},
  author =       {Mignacco, Francesca and Krzakala, Florent and Lu, Yue and Urbani, Pierfrancesco and Zdeborova, Lenka},
  booktitle = 	 {Proceedings of the 37th International Conference on Machine Learning},
  pages = 	 {6874--6883},
  year = 	 {2020},
  editor = 	 {III, Hal Daumé and Singh, Aarti},
  volume = 	 {119},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {13--18 Jul},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v119/mignacco20a/mignacco20a.pdf},
  url = 	 {https://proceedings.mlr.press/v119/mignacco20a.html},
  abstract = 	 {We consider a high-dimensional mixture of two Gaussians in the noisy regime where even an oracle knowing the centers of the clusters misclassifies a small but finite fraction of the points. We provide a rigorous analysis of the generalization error of regularized convex classifiers, including ridge, hinge and logistic regression, in the high-dimensional limit where the number $n$ of samples and their dimension $d$ go to infinity while their ratio is fixed to $\alpha=n/d$. We discuss surprising effects of the regularization that in some cases allows to reach the Bayes-optimal performances. We also illustrate the interpolation peak at low regularization, and analyze the role of the respective sizes of the two clusters.}
}
The Role of Regularization in Classification of High-dimensional Noisy Gaussian Mixture · ICML 2020