← Search

Massimiliano Pontil

69 accepted papers

2026

Generalization of Gibbs and Langevin Monte Carlo Algorithms in the Interpolation Regime

ICML 2026poster

This paper provides data-dependent bounds on the expected error of the Gibbs algorithm in the overparameterized interpolation regime, where low training errors are also obtained for impossible data, such as random labels in classification. The results show that generalization in the low-temperature …

Cited by 0SourceScholar
2026

Outcome-Aware Spectral Feature Learning for Instrumental Variable Regression

ICML 2026poster

We address the problem of causal effect estimation in the presence of hidden confounders using nonparametric instrumental variable (IV) regression. An established approach is to use estimators based on learned \emph{spectral features}, that is, features spanning the top singular subspaces of the ope…

Cited by 0SourceScholar
2026

Representation Learning for Equivariant Inference with Guarantees

ICML 2026poster

In many real-world applications of regression, conditional probability estimation, and uncertainty quantification, exploiting symmetries rooted in physics or geometry can dramatically improve generalization and sample efficiency. While geometric deep learning has made empirical advances by incorpora…

Cited by 0SourceScholar
2026

Self-Supervised Evolution Operator Learning for High-Dimensional Dynamical Systems

ICLR 2026poster

We introduce an end-to-end approach to learn the evolution operators of large-scale non-linear dynamical systems, such as those describing complex natural phenomena. Evolution operators are particularly well-suited for analyzing systems that exhibit spatio-temporal patterns and have become a key ana…

Cited by 0SourcecodeScholar
2026

Toward Scalable and Valid Conditional Independence Testing with Spectral Representations

ICML 2026poster

Conditional independence (CI) is central to causal inference, feature selection, and graphical modeling, yet it is untestable in many settings without additional assumptions. Existing CI tests often rely on restrictive structural conditions, limiting their validity. Kernel methods using partial cova…

Cited by 0SourceScholar
2025

A Bregman Proximal Viewpoint on Neural Operators

ICML 2025poster

We present several advances on neural operators by viewing the action of operator layers as the minimizers of Bregman regularized optimization problems over Banach function spaces. The proposed framework allows interpreting the activation operators as Bregman proximity operators from dual to primal…

Cited by 0SourcePDFScholar
2025

An Empirical Bernstein Inequality for Dependent Data in Hilbert Spaces and Applications

AISTATS 2025poster

Learning from non-independent and non-identically distributed data poses a persistent challenge in statistical learning. In this study, we introduce data-dependent Bernstein inequalities tailored for vector-valued processes in Hilbert space. Our inequalities apply to both stationary and non-station…

Cited by 0SourceScholar
2025

DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products

NeurIPS 2025poster

Linear Recurrent Neural Networks (linear RNNs) have emerged as competitive alternatives to Transformers for sequence modeling, offering efficient training and linear-time inference. However, existing architectures face a fundamental trade-off between expressivity and efficiency, dictated by the stru…

Cited by 0SourcecodeScholar
2025

Laplace Transform Based Low-Complexity Learning of Continuous Markov Semigroups

ICML 2025poster

Markov processes serve as universal models for many real-world random processes. This paper presents a data-driven approach to learning these models through the spectral decomposition of the infinitesimal generator (IG) of the Markov semigroup. Its unbounded nature complicates traditional methods s…

Cited by 0SourcePDFScholar
2025

Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues

ICLR 2025oral

Linear Recurrent Neural Networks (LRNNs) such as Mamba, RWKV, GLA, mLSTM, and DeltaNet have emerged as efficient alternatives to Transformers for long sequences. However, both Transformers and LRNNs struggle to perform state-tracking, which may impair performance in tasks such as code evaluation. In…

2024

Consistent Long-Term Forecasting of Ergodic Dynamical Systems

ICML 2024poster

We study the problem of forecasting the evolution of a function of the state (observable) of a discrete ergodic dynamical system over multiple time steps. The elegant theory of Koopman and transfer operators can be used to evolve any such function forward in time. However, their estimators are usual…

Cited by 3SourcePDFScholar
2024

From Biased to Unbiased Dynamics: An Infinitesimal Generator Approach

NeurIPS 2024poster

We investigate learning the eigenfunctions of evolution operators for time-reversal invariant stochastic processes, a prime example being the Langevin equation used in molecular dynamics. Many physical or chemical processes described by this equation involve transitions between metastable states sep…

2024

Learning invariant representations of time-homogeneous stochastic dynamical systems

ICLR 2024poster

We consider the general class of time-homogeneous stochastic dynamical systems, both discrete and continuous, and study the problem of learning a representation of the state that faithfully captures its dynamics. This is instrumental to learning the transfer operator or the generator of the system,…

2024

Learning the Infinitesimal Generator of Stochastic Diffusion Processes

NeurIPS 2024poster

We address data-driven learning of the infinitesimal generator of stochastic diffusion processes, essential for understanding numerical simulations of natural and physical systems. The unbounded nature of the generator poses significant challenges, rendering conventional analysis techniques for Hilb…

Cited by 4SourcePDFScholar
2024

Leveraging Symmetry in RL-based Legged Locomotion Control

IROS 2024poster

Model-free reinforcement learning is a promising approach for autonomously solving challenging robotics control problems, but faces exploration difficulty without information about the robot’s morphology. The under-exploration of multiple modalities with symmetric states leads to behaviors that are…

Cited by 10SourceScholar
2024

Neural Conditional Probability for Uncertainty Quantification

NeurIPS 2024poster

We introduce Neural Conditional Probability (NCP), an operator-theoretic approach to learning conditional distributions with a focus on statistical inference tasks. NCP can be used to build conditional confidence regions and extract key statistics such as conditional quantiles, mean, and covarianc…

Cited by 0SourcePDFScholar
2024

Nonsmooth Implicit Differentiation: Deterministic and Stochastic Convergence Rates

ICML 2024poster

We study the problem of efficiently computing the derivative of the fixed-point of a parametric nondifferentiable contraction map. This problem has wide applications in machine learning, including hyperparameter optimization, meta-learning and data poisoning attacks. We analyze two popular approache…

2024

Operator World Models for Reinforcement Learning

NeurIPS 2024poster

Policy Mirror Descent (PMD) is a powerful and theoretically sound methodology for sequential decision-making. However, it is not directly applicable to Reinforcement Learning (RL) due to the inaccessibility of explicit action-value functions. We address this challenge by introducing a novel approach…

2023

Estimating Koopman operators with sketching to provably learn large scale dynamical systems

NeurIPS 2023poster

The theory of Koopman operators allows to deploy non-parametric machine learning algorithms to predict and analyze complex dynamical systems. Estimators such as principal component regression (PCR) or reduced rank regression (RRR) in kernel spaces can be shown to provably learn Koopman operators fro…

2023

Multi-task Representation Learning with Stochastic Linear Bandits

AISTATS 2023poster

We study the problem of transfer-learning in the setting of stochastic linear contextual bandit tasks. We consider that a low dimensional linear representation is shared across the tasks, and study the benefit of learning the tasks jointly. Following recent results to design Lasso stochastic bandit…

Cited by 28SourcePDFScholar
2023

Sharp Spectral Rates for Koopman Operator Learning

NeurIPS 2023spotlight

Non-linear dynamical systems can be handily described by the associated Koopman operator, whose action evolves every observable of the system forward in time. Learning the Koopman operator and its spectral decomposition from data is enabled by a number of algorithms. In this work we present for the…

2023

Transfer learning for atomistic simulations using GNNs and kernel mean embeddings

NeurIPS 2023poster

Interatomic potentials learned using machine learning methods have been successfully applied to atomistic simulations. However, accurate models require large training datasets, while generating reference calculations is computationally demanding. To bypass this difficulty, we propose a transfer lea…

2022

A gradient estimator via L1-randomization for online zero-order optimization with two point feedback

NeurIPS 2022accept

This work studies online zero-order optimization of convex and Lipschitz functions. We present a novel gradient estimator based on two function evaluations and randomization on the $\ell_1$-sphere. Considering different geometries of feasible sets and Lipschitz assumptions we analyse online dual av…

Cited by 30SourcePDFScholar
2022

Batch Greenkhorn Algorithm for Entropic-Regularized Multimarginal Optimal Transport: Linear Rate of Convergence and Iteration Complexity

ICML 2022spotlight

In this work we propose a batch multimarginal version of the Greenkhorn algorithm for the entropic-regularized optimal transport problem. This framework is general enough to cover, as particular cases, existing Sinkhorn and Greenkhorn algorithms for the bi-marginal setting, and greedy MultiSinkhorn…

Cited by 3SourcePDFScholar
2022

Distribution Regression with Sliced Wasserstein Kernels

ICML 2022spotlight

The problem of learning functions over spaces of probabilities - or distribution regression - is gaining significant interest in the machine learning community. The main challenge in these settings is to identify a suitable representation capturing all relevant properties of a distribution. The well…

2022

Group Meritocratic Fairness in Linear Contextual Bandits

NeurIPS 2022accept

We study the linear contextual bandit problem where an agent has to select one candidate from a pool and each candidate belongs to a sensitive group. In this setting, candidates' rewards may not be directly comparable between groups, for example when the agent is an employer hiring candidates from d…

2022

Implicit kernel meta-learning using kernel integral forms

UAI 2022poster

Meta-learning algorithms have made significant progress in the context of meta-learning for image classification but less attention has been given to the regression setting. In this paper we propose to learn the probability distribution representing a random feature kernel that we wish to use within…

2022

Learning Dynamical Systems via Koopman Operator Regression in Reproducing Kernel Hilbert Spaces

NeurIPS 2022accept

We study a class of dynamical systems modelled as stationary Markov chains that admit an invariant distribution via the corresponding transfer or Koopman operator. While data-driven algorithms to reconstruct such operators are well known, their relationship with statistical learning is largely unexp…

2022

Multi-source domain adaptation via weighted joint distributions optimal transport

UAI 2022poster

This work addresses the problem of domain adaptation on an unlabeled target dataset using knowledge from multiple labelled source datasets. Most current approaches tackle this problem by searching for an embedding that is invariant across source and target domains, which corresponds to searching for…

Cited by 47SourcePDFScholar
2021

Concentration inequalities under sub-Gaussian and sub-exponential conditions

NeurIPS 2021poster

We prove analogues of the popular bounded difference inequality (also called McDiarmid's inequality) for functions of independent random variables under sub-gaussian and sub-exponential conditions. Applied to vector-valued concentration and the method of Rademacher complexities these inequalities al…

Cited by 38SourcePDFScholar
2021

Distance-Based Regularisation of Deep Networks for Fine-Tuning

ICLR 2021poster

We investigate approaches to regularisation during fine-tuning of deep neural networks. First we provide a neural network generalisation bound based on Rademacher complexity that uses the distance the weights have moved from their initial values. This bound has no direct dependence on the number of…

2021

Distributed Zero-Order Optimization under Adversarial Noise

NeurIPS 2021poster

We study the problem of distributed zero-order optimization for a class of strongly convex functions. They are formed by the average of local objectives, associated to different nodes in a prescribed network. We propose a distributed zero-order projected gradient descent algorithm to solve the probl…

Cited by 24SourcePDFScholar
2021

Robust Unsupervised Learning via L-statistic Minimization

ICML 2021spotlight

Designing learning algorithms that are resistant to perturbations of the underlying data distribution is a problem of wide practical and theoretical importance. We present a general approach to this problem focusing on unsupervised learning. The key assumption is that the perturbing distribution is…

Cited by 14SourcePDFScholar
2021

The Role of Global Labels in Few-Shot Classification and How to Infer Them

NeurIPS 2021poster

Few-shot learning is a central problem in meta-learning, where learners must quickly adapt to new tasks given limited training data. Recently, feature pre-training has become a ubiquitous component in state-of-the-art meta-learning methods and is shown to provide significant performance improvement.…

Cited by 18SourcePDFScholar
2020

Exploiting Higher Order Smoothness in Derivative-free Optimization and Continuous Bandits

NeurIPS 2020poster

We address the problem of zero-order optimization of a strongly convex function. The goal is to find the minimizer of the function by a sequential exploration of its function values, under measurement noise. We study the impact of higher order smoothness properties of the function on the optimizatio…

Cited by 60SourcePDFScholar
2020

Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation Learning

NeurIPS 2020poster

Developing learning methods which do not discriminate subgroups in the population is a central goal of algorithmic fairness. One way to reach this goal is by modifying the data representation in order to meet certain fairness constraints. In this work we measure fairness according to demographic par…

2020

Fair regression via plug-in estimator and recalibration with statistical guarantees

NeurIPS 2020oral

We study the problem of learning an optimal regression function subject to a fairness constraint. It requires that, conditionally on the sensitive feature, the distribution of the function output remains the same. This constraint naturally extends the notion of demographic parity, often used in clas…

2020

Fair regression with Wasserstein barycenters

NeurIPS 2020poster

We study the problem of learning a real-valued function that satisfies the Demographic Parity constraint. It demands the distribution of the predicted output to be independent of the sensitive attribute. We consider the case that the sensitive attribute is available for prediction. We establish a co…

2020

Marthe: Scheduling the Learning Rate Via Online Hypergradients

IJCAI 2020poster

We study the problem of fitting task-specific learning rate schedules from the perspective of hyperparameter optimization, aiming at good generalization. We describe the structure of the gradient of a validation error w.r.t. the learning rate schedule -- the hypergradient. Based on this, we introduc…

2020

On the Iteration Complexity of Hypergradient Computation

ICML 2020poster

We study a general class of bilevel problems, consisting in the minimization of an upper-level objective which depends on the solution to a parametric fixed-point equation. Important instances arising in machine learning include hyperparameter optimization, meta-learning, and certain graph and recur…

2020

Online Parameter-Free Learning of Multiple Low Variance Tasks

UAI 2020poster

We propose a method to learn a common bias vector for a growing sequence of low-variance tasks. Unlike state-of-the-art approaches, our method does not require tuning any hyper-parameter. Our approach is presented in the non-statistical setting and can be of two variants. The “aggressive” one update…

2020

The Advantage of Conditional Meta-Learning for Biased Regularization and Fine Tuning

NeurIPS 2020poster

Biased regularization and fine tuning are two recent meta-learning approaches. They have been shown to be effective to tackle distributions of tasks, in which the tasks’ target vectors are all close to a common meta-parameter vector. However, these methods may perform poorly on heterogeneous environ…

2019

Fast and Continuous Foothold Adaptation for Dynamic Locomotion Through CNNs

RA-L 2019

Legged robots can outperform wheeled machines for most navigation tasks across unknown and rough terrains. For such tasks, visual feedback is a fundamental asset to provide robots with terrain awareness. However, robust dynamic locomotion on difficult terrains with real-time performance guarantees r

Cited by 81SourceScholar
2019

Learning Discrete Structures for Graph Neural Networks

ICML 2019oral

Graph neural networks (GNNs) are a popular class of machine learning models that have been successfully applied to a range of problems. Their major advantage lies in their ability to explicitly incorporate a sparse and discrete dependency structure between data points. Unfortunately, GNNs can only b…

Cited by 526SourcePDFScholar
2019

Learning-to-Learn Stochastic Gradient Descent with Biased Regularization

ICML 2019oral

We study the problem of learning-to-learn: infer- ring a learning algorithm that works well on a family of tasks sampled from an unknown distribution. As class of algorithms we consider Stochastic Gradient Descent (SGD) on the true risk regularized by the square euclidean distance from a bias vector…

2019

Leveraging Labeled and Unlabeled Data for Consistent Fair Binary Classification

NeurIPS 2019poster

We study the problem of fair binary classification using the notion of Equal Opportunity. It requires the true positive rate to distribute equally across the sensitive groups. Within this setting we show that the fair optimal classifier is obtained by recalibrating the Bayes classifier by a group-de…

2019

Leveraging Low-Rank Relations Between Surrogate Tasks in Structured Prediction

ICML 2019oral

We study the interplay between surrogate methods for structured prediction and techniques from multitask learning designed to leverage relationships between surrogate outputs. We propose an efficient algorithm based on trace norm regularization which, differently from previous methods, does not requ…

2019

Online-Within-Online Meta-Learning

NeurIPS 2019poster

We study the problem of learning a series of tasks in a fully online Meta-Learning setting. The goal is to exploit similarities among the tasks to incrementally adapt an inner online algorithm in order to incur a low averaged cumulative error over the tasks. We focus on a family of inner algorithms…

2019

Sinkhorn Barycenters with Free Support via Frank-Wolfe Algorithm

NeurIPS 2019spotlight

We present a novel algorithm to estimate the barycenter of arbitrary probability distributions with respect to the Sinkhorn divergence. Based on a Frank-Wolfe optimization strategy, our approach proceeds by populating the support of the barycenter incrementally, without requiring any pre-allocation.…

2018

Bilevel Programming for Hyperparameter Optimization and Meta-Learning

ICML 2018oral

We introduce a framework based on bilevel programming that unifies gradient-based hyperparameter optimization and meta-learning. We show that an approximate version of the bilevel problem can be solved by taking into explicit account the optimization dynamics for the inner objective. Depending on th…

Cited by 932SourcePDFScholar
2018

Differential Properties of Sinkhorn Approximation for Learning with Wasserstein Distance

NeurIPS 2018poster

Applications of optimal transport have recently gained remarkable attention as a result of the computational advantages of entropic regularization. However, in most situations the Sinkhorn approximation to the Wasserstein distance is replaced by a regularized version that is less accurate but easy…

2018

Empirical Risk Minimization Under Fairness Constraints

NeurIPS 2018poster

We address the problem of algorithmic fairness: ensuring that sensitive information does not unfairly influence the outcome of a classifier. We present an approach based on empirical risk minimization, which incorporates a fairness constraint into the learning problem. It encourages the conditional…

2017

Consistent Multitask Learning with Nonlinear Output Relations

NeurIPS 2017poster

Key to multitask learning is exploiting the relationships between different tasks to improve prediction performance. Most previous methods have focused on the case where tasks relations can be modeled as linear operators and regularization approaches can be used successfully. However, in practice as…

Cited by 40SourcePDFScholar
2017

Forward and Reverse Gradient-Based Hyperparameter Optimization

ICML 2017poster

We study two procedures (reverse-mode and forward-mode) for computing the gradient of the validation error with respect to the hyperparameters of any iterative learning algorithm such as stochastic gradient descent. These procedures mirror two ways of computing gradients for recurrent neural network…

Cited by 569SourcePDFScholar
2016

Unsupervised Cross-Dataset Transfer Learning for Person Re-Identification

CVPR 2016poster

Most existing person re-identification (Re-ID) approaches follow a supervised learning framework, in which a large number of labelled matching pairs are required for training. This severely limits their scalability in real-world applications. To overcome this limitation, we develop a novel cross-dat…

Cited by 457PDFScholar
2015

Learning With Dataset Bias in Latent Subcategory Models

CVPR 2015poster

Latent subcategory models (LSMs) offer significant improvements over training flat classifiers such as linear SVMs. Training LSMs is a challenging task due to the potentially large number of local optima in the objective function and the increased model complexity which requires large training set s…

Cited by 17SourcePDFScholar