← Search

Romain COUILLET

35 accepted papers

2023

Asymptotic Bayes risk of semi-supervised multitask learning on Gaussian mixture

AISTATS 2023poster

The article considers semi-supervised multitask learning on a Gaussian mixture model (GMM). Using methods from statistical physics, we compute the asymptotic Bayes risk of each task in the regime of large datasets in high dimension, from which we analyze the role of task similarity in learning and e…

2023

Large Dimensional Analysis of LS-SVM Transfer Learning: Application to Polsar Classification

ICASSP 2023accepted

This article analyzes a kernel-based transfer learning method, under a k-class Gaussian mixture model for the input data. Following recent advances in random matrix theory, we propose new insights in transfer learning schemes for challenging cases, when the first-order statistics of all data classes…

Cited by 0SourceScholar
2022

A Random Matrix Analysis of Data Stream Clustering: Coping With Limited Memory Resources

ICML 2022spotlight

This article introduces a random matrix framework for the analysis of clustering on high-dimensional data streams, a particularly relevant setting for a more sober processing of large amounts of data with limited memory and energy resources. Assuming data $\mathbf{x}_1, \mathbf{x}_2, \ldots$ arrives…

Cited by 6SourcePDFScholar
2022

Random matrices in service of ML footprint: ternary random features with no performance loss

ICLR 2022poster

In this article, we investigate the spectral behavior of random features kernel matrices of the type ${\bf K} = \mathbb{E}_{{\bf w}} \left[\sigma\left({\bf w}^{\sf T}{\bf x}_i\right)\sigma\left({\bf w}^{\sf T}{\bf x}_j\right)\right]_{i,j=1}^n$, with nonlinear function $\sigma(\cdot)$, data ${\bf x}_…

2021

Deciphering and Optimizing Multi-Task Learning: a Random Matrix Approach

ICLR 2021spotlight

This article provides theoretical insights into the inner workings of multi-task and transfer learning methods, by studying the tractable least-square support vector machine multi-task learning (LS-SVM MTL) method, in the limit of large ($p$) and numerous ($n$) data. By a random matrix analysis appl…

Cited by 11SourcePDFScholar
2021

The Unexpected Deterministic and Universal Behavior of Large Softmax Classifiers

AISTATS 2021poster

This paper provides a large dimensional analysis of the Softmax classifier. We discover and prove that, when the classifier is trained on data satisfying loose statistical modeling assumptions, its weights become deterministic and solely depend on the data statistical means and covariances. As a str…

2021

Two-way kernel matrix puncturing: towards resource-efficient PCA and spectral clustering

ICML 2021spotlight

The article introduces an elementary cost and storage reduction method for spectral clustering and principal component analysis. The method consists in randomly “puncturing” both the data matrix $X\in\mathbb{C}^{p\times n}$ (or $\mathbb{R}^{p\times n}$) and its corresponding kernel (Gram) matrix $K$…

Cited by 17SourcePDFScholar
2020

A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent

NeurIPS 2020poster

This article characterizes the exact asymptotics of random Fourier feature (RFF) regression, in the realistic setting where the number of data samples $n$, their dimension $p$, and the dimension of feature space $N$ are all large and comparable. In this regime, the random RFF Gram matrix no longer c…

Cited by 133SourcePDFScholar
2020

Community detection in sparse time-evolving graphs with a dynamical Bethe-Hessian

NeurIPS 2020poster

This article considers the problem of community detection in sparse dynamical graphs in which the community structure evolves over time. A fast spectral algorithm based on an extension of the Bethe-Hessian matrix is proposed, which benefits from the positive correlation in the class labels and in th…

2020

Optimal Laplacian Regularization for Sparse Spectral Community Detection

ICASSP 2020accepted

Regularization of the classical Laplacian matrices was empirically shown to improve spectral clustering in sparse networks. It was observed that small regularizations are preferable, but this point was left as a heuristic argument. In this paper we formally determine a proper regularization which is…

Cited by 0SourceScholar
2020

Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian Mixtures

ICML 2020poster

This paper shows that deep learning (DL) representations of data produced by generative adversarial nets (GANs) are random vectors which fall within the class of so-called \emph{concentrated} random vectors. Further exploiting the fact that Gram matrices, of the type $G = X^\intercal X$ with $X=[x_1…

Cited by 85SourcePDFScholar
2019

A Kernel Random Matrix-Based Approach for Sparse PCA

ICLR 2019poster

In this paper, we present a random matrix approach to recover sparse principal components from n p-dimensional vectors. Specifically, considering the large dimensional setting where n, p → ∞ with p/n → c ∈ (0, ∞) and under Gaussian vector observations, we study kernel random matrices of the type f (…

Cited by 17SourcePDFScholar
2019

A Large Scale Analysis of Logistic Regression: Asymptotic Performance and New Insights

ICASSP 2019accepted

Logistic regression, one of the most popular machine learning binary classification methods, has been long believed to be unbiased. In this paper, we consider the "hard" classification problem of separating high dimensional Gaussian vectors, where the data dimension p and the sample size n are both…

Cited by 0SourceScholar
2019

Improved Estimation of the Distance between Covariance Matrices

ICASSP 2019accepted

A wide range of machine learning and signal processing applications involve data discrimination through covariance matrices. A broad family of metrics, among which the Frobe-nius, Fisher, Bhattacharyya distances, as well as the Kullback-Leibler or Rényi divergences, are regularly exploited. Not bein…

Cited by 0SourceScholar
2019

Kernel Random Matrices of Large Concentrated Data: the Example of GAN-Generated Images

ICASSP 2019accepted

Based on recent random matrix advances in the analysis of kernel methods for classification and clustering, this paper proposes the study of large kernel methods for a wide class of random inputs, i.e., concentrated data, which are more generic than Gaussian mixtures. The concentration assumption is…

Cited by 0SourceScholar
2019

Latent Heterogeneous Multilayer Community Detection

ICASSP 2019accepted

We propose a method for simultaneously detecting shared and unshared communities in heterogeneous multilayer weighted and undirected networks. The multilayer network is assumed to follow a generative probabilistic model that takes into account the similarities and dissimilarities between the communi…

Cited by 0SourceScholar
2019

Random Matrix Improved Covariance Estimation for a Large Class of Metrics

ICML 2019oral

Relying on recent advances in statistical estimation of covariance distances based on random matrix theory, this article proposes an improved covariance and precision matrix estimation for a wide family of metrics. The method is shown to largely outperform the sample covariance matrix estimate and t…

Cited by 18SourcePDFScholar
2018

Random Matrix Asymptotics of Inner Product Kernel Spectral Clustering

ICASSP 2018accepted

We study in this article the asymptotic performance of spectral clustering with inner product kernel for Gaussian mixture models of high dimension with numerous samples. As is now classical in large dimensional spectral analysis, we establish a phase transition phenomenon by which a minimum distance…

Cited by 0SourceScholar
2017

The counterintuitive mechanism of graph-based semi-supervised learning in the big data regime

ICASSP 2017accepted

In this article, a new approach is proposed to study the performance of graph-based semi-supervised learning methods, under the assumptions that the dimension of data p and their number n grow large at the same rate and that the data arise from a Gaussian mixture model. Unlike small dimensional syst…

Cited by 0SourceScholar
2016

A Random Matrix Approach to Echo-State Neural Networks

ICML 2016poster

Recurrent neural networks, especially in their linear version, have provided many qualitative insights on their performance under different configurations. This article provides, through a novel random matrix framework, the quantitative counterpart of these performance results, specifically in the c…

Cited by 7SourcePDFScholar
2015

Large dimensional analysis of Maronna's M-estimator with outliers

ICASSP 2015accepted

Building on recent results in the random matrix analysis of robust estimators of scatter, we show that a certain class of such estimators obtained from samples containing outliers behaves similar to a well-known random matrix model in the limiting regime where both the population and sample sizes gr…

Cited by 0SourceScholar
2015

Second order statistics of bilinear forms of robust scatter estimators

ICASSP 2015accepted

This paper lies in the lineage of recent works studying the asymptotic behaviour of robust-scatter estimators in the case where the number of observations and the dimension of the population covariance matrix grow at infinity with the same pace. In particular, we analyze the fluctuations of bilinear…

Cited by 0SourceScholar