← Search

Joseph Salmon

23 accepted papers

2025

Class conditional conformal prediction for multiple inputs by p-value aggregation

NeurIPS 2025poster

Conformal prediction methods are statistical tools designed to quantify uncertainty and generate predictive sets with guaranteed coverage probabilities. This work introduces an innovative refinement to these methods for classification tasks, specifically tailored for scenarios where multiple observa…

Cited by 0SourceScholar
2023

High-Dimensional Private Empirical Risk Minimization by Greedy Coordinate Descent

AISTATS 2023poster

In this paper, we study differentially private empirical risk minimization (DP-ERM). It has been shown that the worst-case utility of DP-ERM reduces polynomially as the dimension increases. This is a major obstacle to privately learning large machine learning models. In high dimension, it is common…

Cited by 7SourcePDFScholar
2022

Benchopt: Reproducible, efficient and collaborative optimization benchmarks

NeurIPS 2022accept

Numerical validation is at the core of machine learning research as it allows us to assess the actual impact of new methods, and to confirm the agreement between theory and practice. Yet, the rapid development of the field poses several challenges: researchers are confronted with a profusion of meth…

2022

Convergent Working Set Algorithm for Lasso with Non-Convex Sparse Regularizers

AISTATS 2022poster

Non-convex sparse regularizers are common tools for learning with high-dimensional data. For accelerating convergence for Lasso problem involving those regularizers, a working set strategy addresses the optimization problem through an iterative algorithm by gradually incrementing the number of varia…

2022

Differentially Private Coordinate Descent for Composite Empirical Risk Minimization

ICML 2022spotlight

Machine learning models can leak information about the data used to train them. To mitigate this issue, Differentially Private (DP) variants of optimization algorithms like Stochastic Gradient Descent (DP-SGD) have been designed to trade-off utility for privacy in Empirical Risk Minimization (ERM) p…

Cited by 21SourcePDFScholar
2022

Stochastic smoothing of the top-K calibrated hinge loss for deep imbalanced classification

ICML 2022spotlight

In modern classification tasks, the number of labels is getting larger and larger, as is the size of the datasets encountered in practice. As the number of classes increases, class ambiguity and class imbalance become more and more problematic to achieve high top-1 accuracy. Meanwhile, Top-K metrics…

2021

Pl@ntNet-300K: a plant image dataset with high label ambiguity and a long-tailed distribution

NeurIPS 2021poster

This paper presents a novel image dataset with high intrinsic ambiguity specifically built for evaluating and comparing set-valued classifiers. This dataset, built from the database of Pl@ntnet citizen observatory, consists of 306,146 images covering 1,081 species. We highlight two particular featur…

Cited by 50SourceScholar
2020

Implicit differentiation of Lasso-type models for hyperparameter optimization

ICML 2020poster

Setting regularization parameters for Lasso-type estimators is notoriously difficult, though crucial for obtaining the best accuracy. The most popular hyperparameter optimization approach is grid-search on a held-out dataset. However, grid-search requires to choose a predefined grid of parameters an…

2020

Statistical control for spatio-temporal MEG/EEG source imaging with desparsified mutli-task Lasso

NeurIPS 2020poster

Detecting where and when brain regions activate in a cognitive task or in a given clinical condition is the promise of non-invasive techniques like magnetoencephalography (MEG) or electroencephalography (EEG). This problem, referred to as source localization, or source imaging, poses however a high-…

2020

Support recovery and sup-norm convergence rates for sparse pivotal estimation

AISTATS 2020poster

In high dimensional sparse regression, pivotal estimators are estimators for which the optimal regularization parameter is independent of the noise level. The canonical pivotal estimator is the square-root Lasso, formulated along with its derivatives as a “non-smooth + non-smooth” optimization probl…

Cited by 8SourcePDFScholar
2019

Handling correlated and repeated measurements with the smoothed multivariate square-root Lasso

NeurIPS 2019poster

A limitation of Lasso-type estimators is that the optimal regularization parameter depends on the unknown noise level. Estimators such as the concomitant Lasso address this dependence by jointly estimating the noise level and the regression coefficients. Additionally, in many applications, the data…

2019

Safe Grid Search with Optimal Complexity

ICML 2019oral

Popular machine learning estimators involve regularization parameters that can be challenging to tune, and standard strategies rely on grid search for this task. In this paper, we revisit the techniques of approximating the regularization path up to predefined tolerance $\epsilon$ in a unified frame…

2019

Screening rules for Lasso with non-convex Sparse Regularizers

ICML 2019oral

Leveraging on the convexity of the Lasso problem, screening rules help in accelerating solvers by discarding irrelevant variables, during the optimization process. However, because they provide better theoretical guarantees in identifying relevant variables, several non-convex regularizers for the L…

2018

Celer: a Fast Solver for the Lasso with Dual Extrapolation

ICML 2018oral

Convex sparsity-inducing regularizations are ubiquitous in high-dimensional machine learning, but solving the resulting optimization problems can be slow. To accelerate solvers, state-of-the-art approaches consist in reducing the size of the optimization problem at hand. In the context of regression…

2018

Generalized Concomitant Multi-Task Lasso for Sparse Multimodal Regression

AISTATS 2018poster

In high dimension, it is customary to consider Lasso-type estimators to enforce sparsity. For standard Lasso theory to hold, the regularization parameter should be proportional to the noise level, which is often unknown in practice. A remedy is to consider estimators such as the Concomitant Lasso, w…

2016

GAP Safe Screening Rules for Sparse-Group Lasso

NeurIPS 2016poster

For statistical learning in high dimension, sparse regularizations have proven useful to boost both computational and statistical efficiency. In some contexts, it is natural to handle more refined structures than pure sparsity, such as for instance group sparsity. Sparse-Group Lasso has recently bee…

2016

Gossip Dual Averaging for Decentralized Optimization of Pairwise Functions

ICML 2016poster

In decentralized networks (of sensors, connected objects, etc.), there is an important need for efficient algorithms to optimize a global cost function, for instance to learn a global model from the local data collected by each computing unit. In this paper, we address the problem of decentralized m…

Cited by 123SourcePDFScholar
2015

Extending Gossip Algorithms to Distributed Estimation of U-statistics

NeurIPS 2015spotlight

Efficient and robust algorithms for decentralized estimation in networks are essential to many distributed systems. Whereas distributed estimation of sample mean statistics has been the subject of a good deal of attention, computation of U-statistics, relying on more expensive averaging over pairs o…

Cited by 15SourcePDFScholar
2015

GAP Safe screening rules for sparse multi-task and multi-class models

NeurIPS 2015poster

High dimensional regression benefits from sparsity promoting regularizations. Screening rules leverage the known sparsity of the solution by ignoring some variables in the optimization, hence speeding up solvers. When the procedure is proven not to discard features wrongly the rules are said to be s…

Cited by 90SourcePDFScholar