← Search

Stephan Mandt

63 accepted papers

2026

Calibrated Test-Time Guidance for Bayesian Inference

ICML 2026poster

Test-time guidance is a widely used mechanism for steering pre-trained diffusion models toward outcomes specified by a reward function. Existing approaches, however, focus on reward maximization rather than sampling from the true Bayesian posterior, leading to miscalibrated inference. In this work, …

Cited by 0SourceScholar
2026

Heavy-tailed Physics-Informed Neural Networks

ICML 2026poster

Physics-informed neural networks (PINNs) enforce physical laws by minimizing partial differential equation (PDE) residuals and auxiliary constraints. Standard training relies on a mean-squared error (MSE) objective, which implicitly assumes independent Gaussian residuals with a fixed global variance…

Cited by 0SourceScholar
2026

Parallel Token Generation for Language Models

ICLR 2026poster

Autoregressive transformers are the backbone of modern large language models. Despite their success, inference remains slow due to strictly sequential prediction. Prior attempts to predict multiple tokens per step typically impose independence assumptions across tokens, which limits their ability to…

Cited by 0SourceScholar
2026

Skipping the Zeros in Diffusion Models for Sparse Data Generation

ICML 2026poster

Diffusion models (DMs) excel on dense continuous data, but are not designed for sparse continuous data. They do not model exact zeros that represent the deliberate absence of a signal. As a result, they erase sparsity patterns and perform unnecessary computation on mostly zero entries. With Sparsity…

Cited by 0SourceScholar
2025

AstroCompress: A benchmark dataset for multi-purpose compression of astronomical data

ICLR 2025poster

The site conditions that make astronomical observatories in space and on the ground so desirable---cold and dark---demand a physical remoteness that leads to limited data transmission capabilities. Such transmission limitations directly bottleneck the amount of data acquired and in an era of costly…

2025

Generative Uncertainty in Diffusion Models

UAI 2025

Diffusion models have recently driven significant breakthroughs in generative modeling. While state-of-the-art models produce high-quality samples on average, individual samples can still be low quality. Detecting such samples without human inspection remains a challenging task. To address this, we

2025

Heavy-Tailed Diffusion Models

ICLR 2025poster

Diffusion models achieve state-of-the-art generation quality across many applications, but their ability to capture rare or extreme events in heavy-tailed distributions remains unclear. In this work, we show that traditional diffusion and flow-matching models with standard Gaussian priors fail to ca…

Cited by 6SourcePDFScholar
2025

NoBOOM: Chemical Process Datasets for Industrial Anomaly Detection

NeurIPS 2025poster

Monitoring chemical processes is essential to prevent catastrophic failures, optimize costs and profits, and ensure the safety of employees and the environment. A key component of modern monitoring systems is the automated detection of anomalies in sensor data over time, called time series, enablin…

Cited by 0SourceScholar
2025

One Diffusion to Generate Them All

CVPR 2025poster

We introduce \texttt OneDiffusion - a single large-scale diffusion model designed to tackle a wide range of image synthesis and understanding tasks. It can generate images conditioned on text, depth, pose, layout, or semantic maps. It also handles super-resolution, multi-view generation, instant p…

2025

Transformers for Mixed-type Event Sequences

NeurIPS 2025spotlight

Event sequences appear widely in domains such as medicine, finance, and remote sensing, yet modeling them is challenging due to their heterogeneity: sequences often contain multiple event types with diverse structures—for example, electronic health records that mix discrete events like medical proce…

Cited by 0SourcecodeScholar
2025

UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis

NeurIPS 2025poster

Novel view synthesis (NVS) seeks to render photorealistic, 3D‑consistent images of a scene from unseen camera poses given only a sparse set of posed views. Existing deterministic networks render observed regions quickly but blur unobserved areas, whereas stochastic diffusion‑based methods hallucinat…

Cited by 0SourceScholar
2025

Variational Control for Guidance in Diffusion Models

ICML 2025poster

Diffusion models exhibit excellent sample quality, but existing guidance methods often require additional model training or are limited to specific tasks. We revisit guidance in diffusion models from the perspective of variational inference and control, introducing \emph{Diffusion Trajectory Matchin…

2024

Early-Exit Neural Networks with Nested Prediction Sets

UAI 2024poster

Early-exit neural networks (EENNs) facilitate adaptive inference by producing predictions at multiple stages of the forward pass. In safety-critical applications, these predictions are only meaningful when complemented with reliable uncertainty estimates. Yet, due to their sequential structure, an…

Cited by 1SourcePDFScholar
2024

Fast samplers for Inverse Problems in Iterative Refinement models

NeurIPS 2024poster

Constructing fast samplers for unconditional diffusion and flow-matching models has received much attention recently; however, existing methods for solving *inverse problems*, such as super-resolution, inpainting, or deblurring, still require hundreds to thousands of iterative steps to obtain high-q…

2024

Neural NeRF Compression

ICML 2024poster

Neural Radiance Fields (NeRFs) have emerged as powerful tools for capturing detailed 3D scenes through continuous volumetric representations. Recent NeRFs utilize feature grids to improve rendering quality and speed; however, these representations introduce significant storage overhead. This paper p…

Cited by 1SourcePDFScholar
2024

Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI

ICML 2024poster

In the current landscape of deep learning research, there is a predominant emphasis on achieving high predictive accuracy in supervised tasks involving large image and language datasets. However, a broader perspective reveals a multitude of overlooked metrics, tasks, and data types, such as uncertai…

Cited by 36SourcePDFScholar
2024

Precipitation Downscaling with Spatiotemporal Video Diffusion

NeurIPS 2024poster

In climate science and meteorology, high-resolution local precipitation (rain and snowfall) predictions are limited by the computational costs of simulation-based methods. Statistical downscaling, or super-resolution, is a common workaround where a low-resolution prediction is improved using statist…

Cited by 5SourcePDFScholar
2024

Understanding Pathologies of Deep Heteroskedastic Regression

UAI 2024poster

Deep, overparameterized regression models are notorious for their tendency to overfit. This problem is exacerbated in heteroskedastic models, which predict both mean and residual noise for each data point. At one extreme, these models fit all training data perfectly, eliminating residual noise entir…

Cited by 5SourcePDFScholar
2024

Unity by Diversity: Improved Representation Learning for Multimodal VAEs

NeurIPS 2024poster

Variational Autoencoders for multimodal data hold promise for many tasks in data analysis, such as representation learning, conditional generation, and imputation. Current architectures either share the encoder output, decoder input, or both across modalities to learn a shared representation. Such…

Cited by 4SourcePDFScholar
2023

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

NeurIPS 2023oral

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of hi…

2023

Deep Anomaly Detection under Labeling Budget Constraints

ICML 2023poster

Selecting informative data points for expert feedback can significantly improve the performance of anomaly detection (AD) in various contexts, such as medical diagnostics or fraud detection. In this paper, we determine a set of theoretical conditions under which anomaly scores generalize from labele…

2023

Estimating the Rate-Distortion Function by Wasserstein Gradient Descent

NeurIPS 2023poster

In the theory of lossy compression, the rate-distortion (R-D) function $R(D)$ describes how much a data source can be compressed (in bit-rate) at any given level of fidelity (distortion). Obtaining $R(D)$ for a given data source establishes the fundamental performance limit for all compression algor…

2023

Fully Bayesian Autoencoders with Latent Sparse Gaussian Processes

ICML 2023poster

We present a fully Bayesian autoencoder model that treats both local latent variables and global decoder parameters in a Bayesian fashion. This approach allows for flexible priors and posterior approximations while keeping the inference costs low. To achieve this, we introduce an amortized MCMC appr…

Cited by 7SourcePDFScholar
2023

Probabilistic Querying of Continuous-Time Event Sequences

AISTATS 2023poster

Continuous-time event sequences, i.e., sequences consisting of continuous time stamps and associated event types (“marks”), are an important type of sequential data with many applications, e.g., in clinical medicine or user behavior modeling. Since these data are typically modeled in an autoregressi…

Cited by 4SourcePDFScholar
2023

Zero-Shot Anomaly Detection via Batch Normalization

NeurIPS 2023poster

Anomaly detection (AD) plays a crucial role in many safety-critical application domains. The challenge of adapting an anomaly detector to drift in the normal data distribution, especially when no training data is available for the "new normal," has led to the development of zero-shot AD techniques.…

2022

Latent Outlier Exposure for Anomaly Detection with Contaminated Data

ICML 2022spotlight

Anomaly detection aims at identifying data points that show systematic deviations from the majority of data in an unlabeled dataset. A common assumption is that clean training data (free of anomalies) is available, which is often violated in practice. We propose a strategy for training an anomaly de…

2022

Predictive Querying for Autoregressive Neural Sequence Models

NeurIPS 2022accept

In reasoning about sequential events it is natural to pose probabilistic queries such as “when will event A occur next” or “what is the probability of A occurring before B”, with applications in areas such as user modeling, language models, medicine, and finance. These types of queries are complex t…

2021

Detecting and Adapting to Irregular Distribution Shifts in Bayesian Online Learning

NeurIPS 2021poster

We consider the problem of online learning in the presence of distribution shifts that occur at an unknown rate and of unknown intensity. We derive a new Bayesian online inference approach to simultaneously infer these distribution shifts and adapt the model to the detected changes by integrating id…

2021

Hierarchical Autoregressive Modeling for Neural Video Compression

ICLR 2021poster

Recent work by Marino et al. (2020) showed improved performance in sequential density estimation by combining masked autoregressive flows with hierarchical latent variable models. We draw a connection between such autoregressive generative models and the task of lossy video compression. Specifically…

Cited by 52SourcePDFScholar
2021

Neural Transformation Learning for Deep Anomaly Detection Beyond Images

ICML 2021spotlight

Data transformations (e.g. rotations, reflections, and cropping) play an important role in self-supervised learning. Typically, images are transformed into different views, and neural networks trained on tasks involving these views produce useful feature representations for downstream tasks, includi…

2021

Scalable Gaussian Process Variational Autoencoders

AISTATS 2021poster

Conventional variational autoencoders fail in modeling correlations between data points due to their use of factorized priors. Amortized Gaussian process inference through GP-VAEs has led to significant improvements in this regard, but is still inhibited by the intrinsic complexity of exact GP infer…

2020

GP-VAE: Deep Probabilistic Time Series Imputation

AISTATS 2020poster

Multivariate time series with missing values are common in areas such as healthcare and finance, and have grown in number and complexity over the years. This raises the question whether deep learning methodologies can outperform classical data imputation methods in this domain. However, naive applic…

2020

How Good is the Bayes Posterior in Deep Neural Networks Really?

ICML 2020poster

During the past five years the Bayesian deep learning community has developed increasingly accurate and efficient approximate inference procedures that allow for Bayesian inference in deep neural networks. However, despite this algorithmic progress and the promise of improved uncertainty quantificat…

2020

The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural Networks

ICML 2020poster

Variational Bayesian Inference is a popular methodology for approximating posterior distributions over Bayesian neural network weights. Recent work developing this class of methods has explored ever richer parameterizations of the approximate posterior in the hope of improving performance. In contra…

Cited by 71SourcePDFScholar
2020

User-Dependent Neural Sequence Models for Continuous-Time Event Data

NeurIPS 2020poster

Continuous-time event data are common in applications such as individual behavior data, financial transactions, and medical health records. Modeling such data can be very challenging, in particular for applications with many different types of events,since it requires a model to predict the event ty…

2018

Learning to Infer

ICLR 2018workshop

Inference models, which replace an optimization-based inference procedure with a learned model, have been fundamental in advancing Bayesian deep learning, the most notable example being variational auto-encoders (VAEs). In this paper, we propose iterative inference models, which learn how to optimiz…

Cited by 7SourceScholar
2018

Scalable Generalized Dynamic Topic Models

AISTATS 2018poster

Dynamic topic models (DTMs) model the evolution of prevalent themes in literature, online media, and other forms of text over time. DTMs assume that word co-occurrence statistics change continuously and therefore impose continuous stochastic process priors on their model parameters. These dynamical…

2017

Factorized Variational Autoencoders for Modeling Audience Reactions to Movies

CVPR 2017poster

Matrix and tensor factorization methods are often used for finding underlying low-dimensional patterns from noisy data. In this paper, we study non-linear tensor factoriza- tion methods based on deep variational autoencoders. Our approach is well-suited for settings where the relationship between th…

Cited by 68PDFScholar