← Search

Francis R. Bach

9 accepted papers

2020

Batch normalization provably avoids ranks collapse for randomly initialised deep networks

NeurIPS 2020poster

Randomly initialized neural networks are known to become harder to train with increasing depth, unless architectural enhancements like residual connections and batch normalization are used. We here investigate this phenomenon by revisiting the connection between random initialization in deep network…

Cited by 73SourcePDFScholar
2020

Dual-Free Stochastic Decentralized Optimization with Variance Reduction

NeurIPS 2020poster

We consider the problem of training machine learning models on distributed data in a decentralized way. For finite-sum problems, fast single-machine algorithms for large datasets rely on stochastic updates combined with variance reduction. Yet, existing decentralized stochastic algorithms either do…

2020

Learning with Differentiable Pertubed Optimizers

NeurIPS 2020poster

Machine learning pipelines often rely on optimizers procedures to make discrete decisions (e.g., sorting, picking closest neighbors, or shortest paths). Although these discrete decisions are easily computed in a forward manner, they break the back-propagation of computational graphs. In order to exp…

Cited by 309SourcePDFScholar
2020

Non-parametric Models for Non-negative Functions

NeurIPS 2020spotlight

Linear models have shown great effectiveness and flexibility in many fields such as machine learning, signal processing and statistics. They can represent rich spaces of functions while preserving the convexity of the optimization problems where they are used, and are simple to evaluate, differenti…

2020

Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear Model

NeurIPS 2020poster

In the context of statistical supervised learning, the noiseless linear model assumes that there exists a deterministic linear relation $Y = \langle \theta_*, \Phi(U) \rangle$ between the random output $Y$ and the random feature vector $\Phi(U)$, a potentially non-linear transformation of the inputs…

Cited by 55SourcePDFScholar
2019

Hyper-parameter Learning for Sparse Structured Probabilistic Models

ICASSP 2019accepted

In this paper, we consider the estimation of hyperparameters for regularization terms commonly used for obtaining structured sparse parameters in signal estimation problems, such as signal denoising. By considering the convex regularization terms as negative log-densities, we propose approximate max…

Cited by 0SourceScholar
2017

Primal-dual algorithms for non-negative matrix factorization with the Kullback-Leibler divergence

ICASSP 2017accepted

Non-negative matrix factorization (NMF) approximates a given matrix as a product of two non-negative matrix factors. Multiplicative algorithms deliver reliable results, but they show slow convergence for high-dimensional data and may be stuck away from local minima. Gradient descent methods have bet…

Cited by 0SourceScholar
2016

A weakly-supervised discriminative model for audio-to-score alignment

ICASSP 2016accepted

In this paper, we consider a new discriminative approach to the problem of audio-to-score alignment. We consider two distinct informations provided by music scores: (i) an exact ordered list of musical events and (ii) an approximate prior information about relative duration of events. We extend the…

Cited by 0SourceScholar
2015

An online EM algorithm in hidden (semi-)Markov models for audio segmentation and clustering

ICASSP 2015accepted

Audio segmentation is an essential problem in many audio signal processing tasks, which tries to segment an audio signal into homogeneous chunks. Rather than separately finding change points and computing similarities between segments, we focus on joint segmentation and clustering, using the framewo…

Cited by 0SourceScholar