← Search

Frank Nielsen

18 accepted papers

2024

A Rate-Distortion View of Uncertainty Quantification

ICML 2024poster

In supervised learning, understanding an input’s proximity to the training data can help a model decide whether it has sufficient evidence for reaching a reliable prediction. While powerful probabilistic models such as Gaussian Processes naturally have this property, deep neural networks often lack…

2024

Hyperbolic Embeddings of Supervised Models

NeurIPS 2024poster

Models of hyperbolic geometry have been successfully used in ML for two main tasks: embedding *models* in unsupervised learning (*e.g.* hierarchies) and embedding *data*. To our knowledge, there are no approaches that provide embeddings for supervised models; even when hyperbolic geometry provides…

Cited by 1SourcePDFScholar
2024

Optimal Transport with Tempered Exponential Measures

AAAI 2024technical

In the field of optimal transport, two prominent subfields face each other: (i) unregularized optimal transport, ``a-la-Kantorovich'', which leads to extremely sparse plans but with algorithms that scale poorly, and (ii) entropic-regularized optimal transport, ``a-la-Sinkhorn-Cuturi'', which gets ne…

Cited by 4SourcePDFScholar
2023

Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep Learning

ICML 2023poster

Riemannian submanifold optimization with momentum is computationally challenging because, to ensure that the iterates remain on the submanifold, we often need to solve difficult differential equations. Here, we simplify such difficulties for a class of structured symmetric positive-definite matrices…

2021

Tractable structured natural-gradient descent using local parameterizations

ICML 2021spotlight

Natural-gradient descent (NGD) on structured parameter spaces (e.g., low-rank covariances) is computationally challenging due to difficult Fisher-matrix computations. We address this issue by using \emph{local-parameter coordinates} to obtain a flexible and efficient NGD method that works well for a…

Cited by 38SourcePDFScholar
2021

q-Paths: Generalizing the geometric annealing path using power means

UAI 2021poster

Many common machine learning methods involve the geometric annealing path, a sequence of intermediate densities between two distributions of interest constructed using the geometric average. While alternatives such as the moment-averaging path have demonstrated performance gains in some settings, th…

2019

Sinkhorn AutoEncoders

UAI 2019poster

Optimal transport offers an alternative to maximum likelihood for learning generative autoencoding models. We show that minimizing the $p$-Wasserstein distance between the generator and the true data distribution is equivalent to the unconstrained min-min optimization of the $p$-Wasserstein distance…

Cited by 127SourcePDFScholar
2017

DeepBach: a Steerable Model for Bach Chorales Generation

ICML 2017poster

This paper introduces DeepBach, a graphical model aimed at modeling polyphonic music and specifically hymn-like pieces. We claim that, after being trained on the chorale harmonizations by Johann Sebastian Bach, our model is capable of generating highly convincing chorales in the style of Bach. DeepB…

2017

Information geometry metric for random signal detection in large random sensing systems

ICASSP 2017accepted

Assume that a N-dimensional noisy measurement vector is available via a N × R linear random sensing operation of a R-dimensional Gaussian signal of interest, denoted by s. The problem statement being addressed here is the study of the minimal Bayes' error probability for the detection of s where N →…

Cited by 0SourceScholar
2016

Comix: Joint estimation and lightspeed comparison of mixture models

ICASSP 2016accepted

The Kullback-Leibler divergence is a widespread dissimilarity measure between probability density functions, based on the Shannon entropy. Unfortunately, there is no analytic formula available to compute this divergence between mixture models, imposing the use of costly approximation algorithms. In…

Cited by 0SourceScholar
2016

Loss factorization, weakly supervised learning and label noise robustness

ICML 2016poster

We prove that the empirical risk of most well-known loss functions factors into a linear term aggregating all labels with a term that is label free, and can further be expressed by sums of the same loss. This holds true even for non-smooth, non-convex losses and in any RKHS. The first term is a (ker…

Cited by 143SourcePDFScholar