← Search

Manfred K Warmuth

13 accepted papers

2024

Hyperbolic Embeddings of Supervised Models

NeurIPS 2024poster

Models of hyperbolic geometry have been successfully used in ML for two main tasks: embedding *models* in unsupervised learning (*e.g.* hierarchies) and embedding *data*. To our knowledge, there are no approaches that provide embeddings for supervised models; even when hyperbolic geometry provides…

Cited by 1SourcePDFScholar
2024

Optimal Transport with Tempered Exponential Measures

AAAI 2024technical

In the field of optimal transport, two prominent subfields face each other: (i) unregularized optimal transport, ``a-la-Kantorovich'', which leads to extremely sparse plans but with algorithms that scale poorly, and (ii) entropic-regularized optimal transport, ``a-la-Sinkhorn-Cuturi'', which gets ne…

Cited by 4SourcePDFScholar
2023

Clustering above Exponential Families with Tempered Exponential Measures

AISTATS 2023poster

The link with exponential families has allowed k-means clustering to be generalized to a wide variety of data-generating distributions in exponential families and clustering distortions among Bregman divergences. Getting the framework to go beyond exponential families is important to lift roadblocks…

Cited by 5SourcePDFScholar
2019

Adaptive Scale-Invariant Online Algorithms for Learning Linear Models

ICML 2019oral

We consider online learning with linear models, where the algorithm predicts on sequentially revealed instances (feature vectors), and is compared against the best linear function (comparator) in hindsight. Popular algorithms in this framework, such as Online Gradient Descent (OGD), have parameters…

Cited by 38SourcePDFScholar
2019

Correcting the bias in least squares regression with volume-rescaled sampling

AISTATS 2019poster

Consider linear regression where the examples are generated by an unknown distribution on R^d x R. Without any assumptions on the noise, the linear least squares solution for any i.i.d. sample will typically be biased w.r.t. the least squares optimum over the entire distribution. However, we show th…

Cited by 18SourcePDFScholar
2019

Robust Bi-Tempered Logistic Loss Based on Bregman Divergences

NeurIPS 2019poster

We introduce a temperature into the exponential function and replace the softmax output layer of the neural networks by a high-temperature generalization. Similarly, the logarithm in the loss we use for training is replaced by a low-temperature logarithm. By tuning the two temperatures, we create lo…

2019

Two-temperature logistic regression based on the Tsallis divergence

AISTATS 2019poster

We develop a variant of multiclass logistic regression that is significantly more robust to noise. The algorithm has one weight vector per class and the surrogate loss is a function of the linear activations (one per class). The surrogate loss of an example with linear activation vector $\mathbf{a}…

Cited by 29SourcePDFScholar