← Search

Xiaoyi Mai

7 accepted papers

2026

Whitening Spherical Gaussian Mixtures in the Large-Dimensional Regime

ICASSP 2026oral

Whitening is a classical technique in unsupervised learning that can facilitate estimation tasks by standardizing data. An important application is the estimation of latent variable models via the decomposition of tensors built from high-order moments. In particular, whitening orthogonalizes the mea…

Cited by 0SourcePDFScholar
2025

The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor Mixtures

ICLR 2025poster

The assumption of Gaussian or Gaussian mixture data has been extensively exploited in a long series of precise performance analyses of machine learning (ML) methods, on large datasets having comparably numerous samples and features. To relax this restrictive assumption, subsequent efforts have been…

Cited by 0SourcePDFScholar
2022

On The Effectiveness of Active Learning by Uncertainty Sampling in Classification of High-Dimensional Gaussian Mixture Data

ICASSP 2022accepted

Active learning aims to reduce the cost of labeling through selective sampling. Despite reported empirical success over passive learning, many popular active learning heuristics such as uncertainty sampling still lack satisfying theoretical guarantees. Towards closing the gap between practical use a…

Cited by 0SourceScholar
2019

A Large Scale Analysis of Logistic Regression: Asymptotic Performance and New Insights

ICASSP 2019accepted

Logistic regression, one of the most popular machine learning binary classification methods, has been long believed to be unbiased. In this paper, we consider the "hard" classification problem of separating high dimensional Gaussian vectors, where the data dimension p and the sample size n are both…

Cited by 0SourceScholar
2017

The counterintuitive mechanism of graph-based semi-supervised learning in the big data regime

ICASSP 2017accepted

In this article, a new approach is proposed to study the performance of graph-based semi-supervised learning methods, under the assumptions that the dimension of data p and their number n grow large at the same rate and that the data arise from a Gaussian mixture model. Unlike small dimensional syst…

Cited by 0SourceScholar