← Search

Rémi Gribonval

31 accepted papers

2026

Path-conditioned training: a principled way to rescale ReLU neural networks

ICML 2026poster

Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network parameters. While two properly rescaled weights implement the same function, the training dynamics can be dramatically different. To offer a fresh perspectiv…

Cited by 0SourceScholar
2025

A Rescaling-Invariant Lipschitz Bound Based on Path-Metrics for Modern ReLU Network Parameterizations

ICML 2025poster

Robustness with respect to weight perturbations underpins guarantees for generalization, pruning and quantization. Existing guarantees rely on *Lipschitz bounds in parameter space*, cover only plain feed-forward MLPs, and break under the ubiquitous neuron-wise rescaling symmetry of ReLU networks. We…

Cited by 0SourcePDFScholar
2025

Transformative or Conservative? Conservation laws for ResNets and Transformers

ICML 2025oral

While conservation laws in gradient flow training dynamics are well understood for (mostly shallow) ReLU and linear networks, their study remains largely unexplored for more practical architectures. For this, we first show that basic building blocks such as ReLU (or linear) shallow networks, with or…

Cited by 0SourcePDFScholar
2024

A path-norm toolkit for modern networks: consequences, promises and challenges

ICLR 2024spotlight

This work introduces the first toolkit around path-norms that fully encompasses general DAG ReLU networks with biases, skip connections and any operation based on the extraction of order statistics: max pooling, GroupSort etc. This toolkit notably allows us to establish generalization bounds for mod…

2024

Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows

ICML 2024poster

Conservation laws are well-established in the context of Euclidean gradient flow dynamics, notably for linear or ReLU neural network training. Yet, their existence and principles for non-Euclidean geometries and momentum-based dynamics remain largely unknown. In this paper, we characterize "all" con…

2023

Abide by the law and follow the flow: conservation laws for gradient flows

NeurIPS 2023oral

Understanding the geometric properties of gradient descent dynamics is a key ingredient in deciphering the recent success of very large machine learning models. A striking observation is that trained over-parameterized models retain some properties of the optimization initialization. This "implicit…

Cited by 12SourcePDFScholar
2023

Does a sparse ReLU network training problem always admit an optimum ?

NeurIPS 2023poster

Given a training set, a loss function, and a neural network architecture, it is often taken for granted that optimal network parameters exist, and a common practice is to apply available optimization algorithms to search for them. In this work, we show that the existence of an optimal solution is no…

Cited by 6SourcePDFScholar
2023

Self-supervised learning with rotation-invariant kernels

ICLR 2023top-25%

We introduce a regularization loss based on kernel mean embeddings with rotation-invariant kernels on the hypersphere (also known as dot-product kernels) for self-supervised learning of image representations. Besides being fully competitive with the state of the art, our method significantly reduces…

2022

Fast Multiscale Diffusion On Graphs

ICASSP 2022accepted

Diffusing a graph signal at multiple scales requires to compute the action of the exponential of as many versions of the Laplacian matrix. Considering the truncated Chebyshev polynomial approximation of the exponential, we derive a tightened bound on the approximation error, allowing thus for a bett…

Cited by 0SourceScholar
2021

Training with Quantization Noise for Extreme Model Compression

ICLR 2021poster

We tackle the problem of producing compact models, maximizing their accuracy for a given model size. A standard solution is to train networks with Quantization Aware Training, where the weights are quantized during training and the gradients approximated with the Straight-Through Estimator. In this…

2020

And the Bit Goes Down: Revisiting the Quantization of Neural Networks

ICLR 2020spotlight

In this paper, we address the problem of reducing the memory footprint of convolutional network architectures. We introduce a vector quantization method that aims at preserving the quality of the reconstruction of the network outputs rather than its weights. The principle of our approach is that it…

Cited by 190SourcecodeScholar
2020

Blaster: An Off-Grid Method for Blind and Regularized Acoustic Echoes Retrieval

ICASSP 2020accepted

Acoustic echoes retrieval is a research topic that is gaining importance in many speech and audio signal processing applications such as speech enhancement, source separation, dereverberation and room geometry estimation. This work proposes a novel approach to blindly retrieve the off-grid timing of…

Cited by 0SourceScholar
2020

Fast Optical System Identification by Numerical Interferometry

ICASSP 2020accepted

We propose a numerical interferometry method for identification of optical multiply-scattering systems when only intensity can be measured. Our method simplifies the calibration of optical transmission matrices from a quadratic to a linear inverse problem by first recovering the phase of the measure…

Cited by 0SourceScholar
2019

Differentially Private Compressive K-means

ICASSP 2019accepted

This work addresses the problem of learning from large collections of data with privacy guarantees. The sketched learning framework proposes to deal with the large scale of datasets by compressing them into a single vector of generalized random moments, from which the learning task is then performed…

Cited by 0SourceScholar
2019

OMP and Continuous Dictionaries: Is k-step Recovery Possible?

ICASSP 2019accepted

In this work, we present new theoretical results on sparse recovery guarantees for a greedy algorithm, orthogonal matching pursuit (OMP), in the context of continuous parametric dictionaries, i.e., made up of an infinite uncountable number of atoms. We build up a family of dictionaries for which k-s…

Cited by 0SourceScholar
2018

Faster and Still Safe: Combining Screening Techniques and Structured Dictionaries to Accelerate the Lasso

ICASSP 2018accepted

Accelerating the solution of the Lasso problem becomes crucial when scaling to very high dimensional data. In this paper, we propose a way to combine two existing acceleration techniques: safe screening tests, which simplify the problem by eliminating useless dictionary atoms; and the use of structu…

Cited by 0SourceScholar
2016

Accelerated spectral clustering using graph filtering of random signals

ICASSP 2016accepted

We build upon recent advances in graph signal processing to propose a faster spectral clustering algorithm. Indeed, classical spectral clustering is based on the computation of the first k eigenvectors of the similarity matrix' Laplacian, whose computation cost, even for sparse matrices, becomes pro…

Cited by 0SourceScholar
2016

Joint estimation of sound source location and boundary impedance with physics-driven cosparse regularization

ICASSP 2016accepted

Indoor acoustic source localization can be efficiently performed by modeling the sound propagation in the room, and by solving the arising inverse problem by means of cosparse regularization and convex optimization techniques. However, previous methods relying on this approach used to assume the kno…

Cited by 0SourceScholar
2016

Membrane shape and boundary conditions estimation using eigenmode decomposition

ICASSP 2016accepted

This paper investigates the problem of estimating the shape or the boundary impedance of a vibrating membrane from acoustic measurements in a limited sub-domain of the membrane. In acoustics, polygonal room shapes are usually estimated through room impulse response measurements. Impedance values of…

Cited by 0SourceScholar
2016

Sketching for large-scale learning of mixture models

ICASSP 2016accepted

Learning parameters from voluminous data can be prohibitive in terms of memory and computational requirements. We propose a "compressive learning" framework where we first sketch the data by computing random generalized moments of the underlying probability distribution, then estimate mixture model…

Cited by 0SourceScholar