← Search

Robert Bamler

14 accepted papers

2025

Well-Defined Function-Space Variational Inference in Bayesian Neural Networks via Regularized KL-Divergence

UAI 2025

Bayesian neural networks (BNN) promise to combine the predictive performance of neural networks with principled uncertainty modeling crucial for safety-critical systems and decision making. However, posterior uncertainties depend on the choice of prior, and finding informative priors in weight-space

2025

Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector

AISTATS 2025poster

We revisit the likelihood ratio between a pretrained large language model (LLM) and its finetuned variant as a criterion for out-of-distribution (OOD) detection. The intuition behind such a criterion is that, the pretrained LLM has the prior knowledge about OOD data due to its large amount of traini…

Cited by 0SourcecodeScholar
2024

Differentiable Annealed Importance Sampling Minimizes The Jensen-Shannon Divergence Between Initial and Target Distribution

ICML 2024poster

Differentiable annealed importance sampling (DAIS), proposed by Geffner & Domke (2021) and Zhang et al. (2021), allows optimizing, among others, over the initial distribution of AIS. In this paper, we show that, in the limit of many transitions, DAIS minimizes the symmetrized KL divergence (Jensen-S…

Cited by 1SourcePDFScholar
2024

FSP-Laplace: Function-Space Priors for the Laplace Approximation in Bayesian Deep Learning

NeurIPS 2024poster

Laplace approximations are popular techniques for endowing deep networks with epistemic uncertainty estimates as they can be applied without altering the predictions of the trained network, and they scale to large models and datasets. While the choice of prior strongly affects the resulting posterio…

Cited by 2SourcePDFScholar
2024

Predictive, scalable and interpretable knowledge tracing on structured domains

ICLR 2024spotlight

Intelligent tutoring systems optimize the selection and timing of learning materials to enhance understanding and long-term retention. This requires estimates of both the learner's progress ("knowledge tracing"; KT), and the prerequisite structure of the learning domain ("knowledge mapping"). While…

2020

User-Dependent Neural Sequence Models for Continuous-Time Event Data

NeurIPS 2020poster

Continuous-time event data are common in applications such as individual behavior data, financial transactions, and medical health records. Modeling such data can be very challenging, in particular for applications with many different types of events,since it requires a model to predict the event ty…