← Search

Umut Simsekli

55 accepted papers

2026

On the Interaction of Compressibility and Adversarial Robustness

ICLR 2026poster

Modern neural networks are expected to simultaneously satisfy a host of desirable properties: accurate fitting to training data, generalization to unseen inputs, parameter and computational efficiency, and robustness to adversarial perturbations. While compressibility and robustness have each been s…

Cited by 0SourceScholar
2026

Tightening the Score Matching Gap for Diffusion Models

ICML 2026poster

Diffusion models (DMs) are a state-of-the-art generative method to approximately sample from an unknown distribution. Their training and evaluation primarily rely on an Evidence Lower Bound (ELBO), which relates the Kullback-Leibler (KL) divergence of model samples to the score matching loss along t…

Cited by 0SourceScholar
2025

Algorithm- and Data-Dependent Generalization Bounds for Diffusion Models

NeurIPS 2025poster

Score-based generative models (SGMs) have emerged as one of the most popular classes of generative models. A substantial body of work now exists on the analysis of SGMs, focusing either on discretization aspects or on their statistical performance. In the latter case, bounds have been derived, under…

Cited by 0SourceScholar
2025

Heavy-Tailed Diffusion with Denoising Levy Probabilistic Models

ICLR 2025poster

Investigating noise distributions beyond Gaussian in diffusion generative models remains an open challenge. The Gaussian case has been a large success experimentally and theoretically, admitting a unified stochastic differential equation (SDE) framework, encompassing score-based and denoising formul…

Cited by 0SourcePDFScholar
2025

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

ICML 2025poster

We show that learning-rate schedules for large model training behave surprisingly similar to a performance bound from non-smooth convex optimization theory. We provide a bound for the constant schedule with linear cooldown; in particular, the practical benefit of cooldown is reflected in the bound d…

2024

Generalization Bounds for Heavy-Tailed SDEs through the Fractional Fokker-Planck Equation

ICML 2024poster

Understanding the generalization properties of heavy-tailed stochastic optimization algorithms has attracted increasing attention over the past years. While illuminating interesting aspects of stochastic optimizers by using heavy-tailed stochastic differential equations as proxies, prior works eithe…

2024

Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD

ICML 2024poster

Neural network compression has been an increasingly important subject, not only due to its practical relevance, but also due to its theoretical implications, as there is an explicit connection between compressibility and generalization error. Recent studies have shown that the choice of the hyperpar…

2024

Piecewise deterministic generative models

NeurIPS 2024poster

We introduce a novel class of generative models based on piecewise deterministic Markov processes (PDMPs), a family of non-diffusive stochastic processes consisting of deterministic motion and random jumps at random times. Similarly to diffusions, such Markov processes admit time reversals that turn…

Cited by 2SourcePDFScholar
2024

Topological Generalization Bounds for Discrete-Time Stochastic Optimization Algorithms

NeurIPS 2024poster

We present a novel set of rigorous and computationally efficient topology-based complexity notions that exhibit a strong correlation with the generalization gap in modern deep neural networks (DNNs). DNNs show remarkable generalization properties, yet the source of these capabilities remains elusive…

Cited by 6SourcePDFScholar
2023

Algorithmic Stability of Heavy-Tailed SGD with General Loss Functions

ICML 2023poster

Heavy-tail phenomena in stochastic gradient descent (SGD) have been reported in several empirical studies. Experimental evidence in previous works suggests a strong interplay between the heaviness of the tails and generalization behavior of SGD. To address this empirical phenomena theoretically, sev…

Cited by 26SourcePDFScholar
2023

Approximate Heavy Tails in Offline (Multi-Pass) Stochastic Gradient Descent

NeurIPS 2023spotlight

A recent line of empirical studies has demonstrated that SGD might exhibit a heavy-tailed behavior in practical settings, and the heaviness of the tails might correlate with the overall performance. In this paper, we investigate the emergence of such heavy tails. Previous works on this problem only…

2023

Efficient Sampling of Stochastic Differential Equations with Positive Semi-Definite Models

NeurIPS 2023poster

This paper deals with the problem of efficient sampling from a stochastic differential equation, given the drift function and the diffusion matrix. The proposed approach leverages a recent model for probabilities (Rudi and Ciliberto, 2021) (the positive semi-definite -- PSD model) from which it is p…

Cited by 1SourcePDFScholar
2023

Generalization Bounds using Data-Dependent Fractal Dimensions

ICML 2023poster

Providing generalization guarantees for modern neural networks has been a crucial task in statistical learning. Recently, several studies have attempted to analyze the generalization error in such settings by using tools from fractal geometry. While these works have successfully introduced new mathe…

Cited by 24SourcePDFScholar
2023

Learning via Wasserstein-Based High Probability Generalisation Bounds

NeurIPS 2023poster

Minimising upper bounds on the population risk or the generalisation gap has been widely used in structural risk minimisation (SRM) -- this is in particular at the core of PAC-Bayesian learning. Despite its successes and unfailing surge of interest in recent years, a limitation of the PAC-Bayesian f…

Cited by 18SourcePDFScholar
2023

Uniform-in-Time Wasserstein Stability Bounds for (Noisy) Stochastic Gradient Descent

NeurIPS 2023poster

Algorithmic stability is an important notion that has proven powerful for deriving generalization bounds for practical algorithms. The last decade has witnessed an increasing number of stability bounds for different algorithms applied on different classes of loss functions. While these bounds have i…

Cited by 9SourcePDFScholar
2022

Chaotic Regularization and Heavy-Tailed Limits for Deterministic Gradient Descent

NeurIPS 2022accept

Recent studies have shown that gradient descent (GD) can achieve improved generalization when its dynamics exhibits a chaotic behavior. However, to obtain the desired effect, the step-size should be chosen sufficiently large, a task which is problem dependent and can be difficult in practice. In thi…

2022

Generalization Bounds for Stochastic Gradient Descent via Localized $\varepsilon$-Covers

NeurIPS 2022accept

In this paper, we propose a new covering technique localized for the trajectories of SGD. This localization provides an algorithm-specific complexity measured by the covering number, which can have dimension-independent cardinality in contrast to standard uniform covering arguments that result in ex…

Cited by 14SourcePDFScholar
2022

Generalization Bounds using Lower Tail Exponents in Stochastic Optimizers

ICML 2022spotlight

Despite the ubiquitous use of stochastic optimization algorithms in machine learning, the precise impact of these algorithms and their dynamics on generalization performance in realistic non-convex settings is still poorly understood. While recent work has revealed connections between generalization…

Cited by 24SourcePDFScholar
2021

Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections

ICML 2021spotlight

Gaussian noise injections (GNIs) are a family of simple and widely-used regularisation methods for training neural networks, where one injects additive or multiplicative Gaussian noise to the network activations at every iteration of the optimisation algorithm, which is typically chosen as stochasti…

2021

Convergence Rates of Stochastic Gradient Descent under Infinite Noise Variance

NeurIPS 2021poster

Recent studies have provided both empirical and theoretical evidence illustrating that heavy tails can emerge in stochastic gradient descent (SGD) in various scenarios. Such heavy tails potentially result in iterates with diverging variance, which hinders the use of conventional convergence analysis…

Cited by 51SourcePDFScholar
2021

Fast Approximation of the Sliced-Wasserstein Distance Using Concentration of Random Projections

NeurIPS 2021poster

The Sliced-Wasserstein distance (SW) is being increasingly used in machine learning applications as an alternative to the Wasserstein distance and offers significant computational and statistical benefits. Since it is defined as an expectation over random projections, SW is commonly approximated by…

2021

Fractal Structure and Generalization Properties of Stochastic Optimization Algorithms

NeurIPS 2021spotlight

Understanding generalization in deep learning has been one of the major challenges in statistical learning theory over the last decade. While recent work has illustrated that the dataset and the training algorithm must be taken into account in order to obtain meaningful generalization bounds, it is…

Cited by 31SourcePDFScholar
2021

Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks

NeurIPS 2021poster

Neural network compression techniques have become increasingly popular as they can drastically reduce the storage and computation requirements for very large networks. Recent empirical studies have illustrated that even simple pruning strategies can be surprisingly effective, and several theoretical…

2021

Intrinsic Dimension, Persistent Homology and Generalization in Neural Networks

NeurIPS 2021poster

Disobeying the classical wisdom of statistical learning theory, modern deep neural networks generalize well even though they typically contain millions of parameters. Recently, it has been shown that the trajectories of iterative optimization algorithms can possess \emph{fractal structures}, and the…

2021

Relative Positional Encoding for Transformers with Linear Complexity

ICML 2021oral

Recent advances in Transformer models allow for unprecedented sequence lengths, due to linear space and time complexity. In the meantime, relative positional encoding (RPE) was proposed as beneficial for classical Transformers and consists in exploiting lags instead of absolute positions for inferen…

2020

Approximate Bayesian Computation with the Sliced-Wasserstein Distance

ICASSP 2020accepted

Approximate Bayesian Computation (ABC) is a popular method for approximate inference in generative models with intractable but easy-to-sample likelihood. It constructs an approximate posterior distribution by finding parameters for which the simulated data are close to the observations in terms of s…

Cited by 0SourceScholar
2020

Explicit Regularisation in Gaussian Noise Injections

NeurIPS 2020poster

We study the regularisation induced in neural networks by Gaussian noise injections (GNIs). Though such injections have been extensively studied when applied to data, there have been few studies on understanding the regularising effect they induce when applied to network activations. Here we derive…

Cited by 78SourcePDFScholar
2020

Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise

ICML 2020poster

Stochastic gradient descent with momentum (SGDm) is one of the most popular optimization algorithms in deep learning. While there is a rich theory of SGDm for convex problems, the theory is considerably less developed in the context of deep learning where the problem is non-convex and the gradient n…

2020

Hausdorff Dimension, Heavy Tails, and Generalization in Neural Networks

NeurIPS 2020spotlight

Despite its success in a wide range of applications, characterizing the generalization properties of stochastic gradient descent (SGD) in non-convex deep learning problems is still an important challenge. While modeling the trajectories of SGD via stochastic differential equations (SDE) under heavy-…

2020

Quantitative Propagation of Chaos for SGD in Wide Neural Networks

NeurIPS 2020poster

In this paper, we investigate the limiting behavior of a continuous-time counterpart of the Stochastic Gradient Descent (SGD) algorithm applied to two-layer overparameterized neural networks, as the number or neurons (i.e., the size of the hidden layer) $N \to \plusinfty$. Following a proba…

Cited by 37SourcePDFScholar
2020

Statistical and Topological Properties of Sliced Probability Divergences

NeurIPS 2020spotlight

The idea of slicing divergences has been proven to be successful when comparing two probability measures in various machine learning applications including generative modeling, and consists in computing the expected value of a `base divergence' between \emph{one-dimensional random projections} of th…

2020

Synchronizing Probability Measures on Rotations via Optimal Transport

CVPR 2020poster

We introduce a new paradigm, `measure synchronization', for synchronizing graphs with measure-valued edges. We formulate this problem as maximization of the cycle-consistency in the space of probability measures over relative rotations. In particular, we aim at estimating marginal distributions of a…

Cited by 37PDFScholar
2019

A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks

ICML 2019oral

The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the classical central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed…

Cited by 289SourcePDFScholar
2019

Asymptotic Guarantees for Learning Generative Models with the Sliced-Wasserstein Distance

NeurIPS 2019spotlight

Minimum expected distance estimation (MEDE) algorithms have been widely used for probabilistic models with intractable likelihood functions and they have become increasingly popular due to their use in implicit generative modeling (e.g.\ Wasserstein generative adversarial networks, Wasserstein autoe…

2019

First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise

NeurIPS 2019poster

Stochastic gradient descent (SGD) has been widely used in machine learning due to its computational efficiency and favorable generalization properties. Recently, it has been empirically demonstrated that the gradient noise in several deep learning settings admits a non-Gaussian, heavy-tailed behavio…

2019

Generalized Sliced Wasserstein Distances

NeurIPS 2019poster

The Wasserstein distance and its variations, e.g., the sliced-Wasserstein (SW) distance, have recently drawn attention from the machine learning community. The SW distance, specifically, was shown to have similar properties to the Wasserstein distance, while being much simpler to compute, and is the…

2019

Non-Asymptotic Analysis of Fractional Langevin Monte Carlo for Non-Convex Optimization

ICML 2019oral

Recent studies on diffusion-based sampling methods have shown that Langevin Monte Carlo (LMC) algorithms can be beneficial for non-convex optimization, and rigorous theoretical guarantees have been proven for both asymptotic and finite-time regimes. Algorithmically, LMC-based algorithms resemble the…

Cited by 36SourcePDFScholar
2019

Probabilistic Permutation Synchronization Using the Riemannian Structure of the Birkhoff Polytope

CVPR 2019oral

We present an entirely new geometric and probabilistic approach to synchronization of correspondences across multiple sets of objects or images. In particular, we present two algorithms: (1) Birkhoff-Riemannian L-BFGS for optimizing the relaxed version of the combinatorially intractable cycle consis…

Cited by 42PDFScholar
2019

Sliced-Wasserstein Flows: Nonparametric Generative Modeling via Optimal Transport and Diffusions

ICML 2019oral

By building upon the recent theory that established the connection between implicit generative modeling (IGM) and optimal transport, in this study, we propose a novel parameter-free algorithm for learning the underlying distributions of complicated datasets and sampling from them. The proposed algor…

2019

Speech Enhancement with Variational Autoencoders and Alpha-stable Distributions

ICASSP 2019accepted

This paper focuses on single-channel semi-supervised speech enhancement. We learn a speaker-independent deep generative speech model using the framework of variational autoencoders. The noise model remains unsupervised because we do not assume prior knowledge of the noisy recording environment. In t…

Cited by 0SourceScholar
2018

Alpha-Stable Low-Rank Plus Residual Decomposition for Speech Enhancement

ICASSP 2018accepted

In this study, we propose a novel probabilistic model for separating clean speech signals from noisy mixtures by decomposing the mixture spectra into a structured speech part and a more flexible residual part. The main novelty in our model is that it uses a family of heavy-tailed distributions, so c…

Cited by 0SourceScholar
2018

Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization

ICML 2018oral

Recent studies have illustrated that stochastic gradient Markov Chain Monte Carlo techniques have a strong potential in non-convex optimization, where local and global convergence guarantees can be shown under certain conditions. By building up on this recent theory, in this study, we develop an asy…

Cited by 27SourcePDFScholar
2018

Bayesian Pose Graph Optimization via Bingham Distributions and Tempered Geodesic MCMC

NeurIPS 2018poster

We introduce Tempered Geodesic Markov Chain Monte Carlo (TG-MCMC) algorithm for initializing pose graph optimization problems, arising in various scenarios such as SFM (structure from motion) or SLAM (simultaneous localization and mapping). TG-MCMC is first of its kind as it unites global non-convex…

Cited by 42SourcePDFScholar
2017

Alpha-stable multichannel audio source separation

ICASSP 2017accepted

In this paper, we focus on modeling multichannel audio signals in the short-time Fourier transform domain for the purpose of source separation. We propose a probabilistic model based on a class of heavy-tailed distributions, in which the observed mixtures and the latent sources are jointly modeled b…

Cited by 0SourceScholar
2017

Learning the Morphology of Brain Signals Using Alpha-Stable Convolutional Sparse Coding

NeurIPS 2017poster

Neural time-series data contain a wide variety of prototypical signal waveforms (atoms) that are of significant importance in clinical and cognitive research. One of the goals for analyzing such data is hence to extract such `shift-invariant' atoms. Even though some success has been reported with ex…

Cited by 63SourcePDFScholar
2017

Parallelized Stochastic Gradient Markov Chain Monte Carlo algorithms for non-negative matrix factorization

ICASSP 2017accepted

Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have become popular in modern data analysis problems due to their computational efficiency. Even though they have proved useful for many statistical models, the application of SG-MCMC to non-negative matrix factorization (NMF) models has…

Cited by 0SourceScholar
2016

Stochastic Gradient Richardson-Romberg Markov Chain Monte Carlo

NeurIPS 2016poster

Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) algorithms have become increasingly popular for Bayesian inference in large-scale applications. Even though these methods have proved useful in several scenarios, their performance is often limited by their bias. In this study, we propose a nove…

Cited by 42SourcePDFScholar
2016

Stochastic thermodynamic integration: Efficient Bayesian model selection via stochastic gradient MCMC

ICASSP 2016accepted

Model selection is a central topic in Bayesian machine learning, which requires the estimation of the marginal likelihood of the data under the models to be compared. During the last decade, conventional model selection methods have lost their charm as they have high computational requirements. In t…

Cited by 0SourceScholar
2015

Learning mixed divergences in coupled matrix and tensor factorization models

ICASSP 2015accepted

Coupled tensor factorization methods are useful for sensor fusion, combining information from several related datasets by simultaneously approximating them by products of latent tensors. In these methods, the choice of a suitable optimization criteria becomes difficult as observed datasets may have…

Cited by 0SourceScholar
2015

Section-level modeling of musical audio for linking performances to scores in Turkish makam music

ICASSP 2015accepted

Section linking aims at relating structural units in the notation of a piece of music to their occurrences in a performance of the piece. In this paper, we address this task by presenting a score-informed hierarchical Hidden Markov Model (HHMM) for modeling musical audio signals on the temporal leve…

Cited by 0SourceScholar