← Search

Yee Whye Teh

88 accepted papers

2026

SigmaDock: Untwisting Molecular Docking with Fragment-Based SE(3) Diffusion

ICLR 2026poster

Determining the binding pose of a ligand to a protein, known as molecular docking, is a fundamental task in drug discovery. Generative approaches promise faster, improved, and more diverse pose sampling than physics-based methods, but are often hindered by chemically implausible outputs, poor genera…

Cited by 0SourcecodeScholar
2026

StochasTok: Improving Fine-Grained Subword Understanding in LLMs

ICLR 2026poster

Subword-level understanding is integral to numerous tasks, including understanding multi-digit numbers, spelling mistakes, abbreviations, rhyming, and wordplay. Despite this, current large language models (LLMs) still struggle disproportionally with seemingly simple subword-level tasks, like countin…

Cited by 0SourcecodeScholar
2025

Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents

ICLR 2025poster

Recent advances in large language models (LLMs) have led to a growing interest in developing LLM-based agents for automating web tasks. However, these agents often struggle with even simple tasks on real-world websites due to their limited capability to understand and process complex web page struct…

2025

Meta-Learning Objectives for Preference Optimization

NeurIPS 2025poster

Evaluating preference optimization (PO) algorithms on LLM alignment is a challenging task that presents prohibitive costs, noise, and several variables like model size and hyper-parameters. In this work, we show that it is possible to gain insights on the efficacy of PO algorithm on much simpler ben…

Cited by 0SourceScholar
2025

SymDiff: Equivariant Diffusion via Stochastic Symmetrisation

ICLR 2025poster

We propose SymDiff, a method for constructing equivariant diffusion models using the framework of stochastic symmetrisation. SymDiff resembles a learned data augmentation that is deployed at sampling time, and is lightweight, computationally efficient, and easy to implement on top of arbitrary off-t…

Cited by 0SourcePDFScholar
2024

Context-Guided Diffusion for Out-of-Distribution Molecular and Protein Design

ICML 2024poster

Generative models have the potential to accelerate key steps in the discovery of novel molecular therapeutics and materials. Diffusion models have recently emerged as a powerful approach, excelling at unconditional sample generation and, with data-driven guidance, conditional generation within their…

2024

EvIL: Evolution Strategies for Generalisable Imitation Learning

ICML 2024poster

Often times in imitation learning (IL), the environment we collect expert demonstrations in and the environment we want to deploy our learned policy in aren't exactly the same (e.g. demonstrations collected in simulation but deployment in the real world). Compared to policy-centric approaches to IL…

2024

Kalman Filter for Online Classification of Non-Stationary Data

ICLR 2024poster

In Online Continual Learning (OCL) a learning system receives a stream of data and sequentially performs prediction and training steps. Key challenges in OCL include automatic adaptation to the specific non-stationary structure of the data and maintaining appropriate predictive uncertainty. To add…

Cited by 8SourcePDFScholar
2024

Non-Stationary Learning of Neural Networks with Automatic Soft Parameter Reset

NeurIPS 2024poster

Neural networks are most often trained under the assumption that data come from a stationary distribution. However, settings in which this assumption is violated are of increasing importance; examples include supervised learning with distributional shifts, reinforcement learning, continual learning…

Cited by 4SourcePDFScholar
2024

Online Adaptation of Language Models with a Memory of Amortized Contexts

NeurIPS 2024poster

Due to the rapid generation and dissemination of information, large language models (LLMs) quickly run out of date despite enormous development costs. To address the crucial need to keep models updated, online learning has emerged as a critical tool when utilizing LLMs for real-world applications. H…

2024

Position: Bayesian Deep Learning is Needed in the Age of Large-Scale AI

ICML 2024poster

In the current landscape of deep learning research, there is a predominant emphasis on achieving high predictive accuracy in supervised tasks involving large image and language datasets. However, a broader perspective reveals a multitude of overlooked metrics, tasks, and data types, such as uncertai…

Cited by 36SourcePDFScholar
2024

SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning

ICLR 2024poster

The recent progress in large language models (LLMs), especially the invention of chain-of-thought prompting, has made it possible to automatically answer questions by stepwise reasoning. However, when faced with more complicated problems that require non-linear thinking, even the strongest LLMs make…

2024

The Edge-of-Reach Problem in Offline Model-Based Reinforcement Learning

NeurIPS 2024poster

Offline reinforcement learning (RL) aims to train agents from pre-collected datasets. However, this comes with the added challenge of estimating the value of behaviors not covered in the dataset. Model-based methods offer a potential solution by training an approximate dynamics model, which then all…

Cited by 1SourcePDFScholar
2024

Unleashing the Power of Meta-tuning for Few-shot Generalization Through Sparse Interpolated Experts

ICML 2024poster

Recent successes suggest that parameter-efficient fine-tuning of foundation models is becoming the state-of-the-art method for transfer learning in vision, gradually replacing the rich literature of alternatives such as meta-learning. In trying to harness the best of both worlds, meta-tuning introdu…

2023

Deep Stochastic Processes via Functional Markov Transition Operators

NeurIPS 2023poster

We introduce Markov Neural Processes (MNPs), a new class of Stochastic Processes (SPs) which are constructed by stacking sequences of neural parameterised Markov transition operators in function space. We prove that these Markov transition operators can preserve the exchangeability and consistency o…

Cited by 7SourcePDFScholar
2023

Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation

ICLR 2023poster

Skip connections and normalisation layers form two standard architectural components that are ubiquitous for the training of Deep Neural Networks (DNNs), but whose precise roles are poorly understood. Recent approaches such as Deep Kernel Shaping have made progress towards reducing our reliance on t…

Cited by 37SourcePDFScholar
2023

Drug Discovery under Covariate Shift with Domain-Informed Prior Distributions over Functions

ICML 2023poster

Accelerating the discovery of novel and more effective therapeutics is an important pharmaceutical problem in which deep learning is playing an increasingly significant role. However, real-world drug discovery tasks are often characterized by a scarcity of labeled data and significant covariate shif…

2023

Geometric Neural Diffusion Processes

NeurIPS 2023poster

Denoising diffusion models have proven to be a flexible and effective paradigm for generative modelling. Their recent extension to infinite dimensional Euclidean spaces has allowed for the modelling of stochastic processes. However, many problems in the natural sciences incorporate symmetries and in…

2023

Learning Instance-Specific Augmentations by Capturing Local Invariances

ICML 2023poster

We introduce InstaAug, a method for automatically learning input-specific augmentations from data. Previous methods for learning augmentations have typically assumed independence between the original input and the transformation applied to that input. This can be highly restrictive, as the invarianc…

2023

Modality-Agnostic Variational Compression of Implicit Neural Representations

ICML 2023poster

We introduce a modality-agnostic neural compression algorithm based on a functional view of data and parameterised as an Implicit Neural Representation (INR). Bridging the gap between latent coding and sparsity, we obtain compact latent representations non-linearly mapped to a soft gating mechanism.…

Cited by 27SourcePDFScholar
2023

Pre-training via Denoising for Molecular Property Prediction

ICLR 2023top-25%

Many important problems involving molecular property prediction from 3D structures have limited data, posing a generalization challenge for neural networks. In this paper, we describe a pre-training technique based on denoising that achieves a new state-of-the-art in molecular property prediction by…

2022

Amortized Rejection Sampling in Universal Probabilistic Programming

AISTATS 2022poster

Naive approaches to amortized inference in probabilistic programs with unbounded loops can produce estimators with infinite variance. This is particularly true of importance sampling inference in programs that explicitly include rejection sampling as part of the user-programmed generative procedure.…

2022

Conformal Off-Policy Prediction in Contextual Bandits

NeurIPS 2022accept

Most off-policy evaluation methods for contextual bandits have focused on the expected outcome of a policy, which is estimated via methods that at best provide only asymptotic guarantees. However, in many applications, the expectation may not be the best measure of performance as it does not capture…

Cited by 20SourcePDFScholar
2022

Continual Learning via Sequential Function-Space Variational Inference

ICML 2022spotlight

Sequential Bayesian inference over predictive functions is a natural framework for continual learning from streams of data. However, applying it to neural networks has proved challenging in practice. Addressing the drawbacks of existing techniques, we propose an optimization objective derived by for…

2022

On Incorporating Inductive Biases into VAEs

ICLR 2022poster

We explain why directly changing the prior can be a surprisingly ineffective mechanism for incorporating inductive biases into variational auto-encoders (VAEs), and introduce a simple and effective alternative approach: Intermediary Latent Space VAEs (InteL-VAEs). InteL-VAEs use an intermediary set…

2022

Riemannian Score-Based Generative Modelling

NeurIPS 2022accept

Score-based generative models (SGMs) are a powerful class of generative models that exhibit remarkable empirical performance. Score-based generative modelling (SGM) consists of a ``noising'' stage, whereby a diffusion is used to gradually add Gaussian noise to data, and a generative model, which ent…

2022

Tractable Function-Space Variational Inference in Bayesian Neural Networks

NeurIPS 2022accept

Reliable predictive uncertainty estimation plays an important role in enabling the deployment of neural networks to safety-critical settings. A popular approach for estimating the predictive uncertainty of neural networks is to define a prior distribution over the network parameters, infer an approx…

2021

BayesIMP: Uncertainty Quantification for Causal Data Fusion

NeurIPS 2021poster

While causal models are becoming one of the mainstays of machine learning, the problem of uncertainty quantification in causal inference remains challenging. In this paper, we study the causal data fusion problem, where data arising from multiple causal graphs are combined to estimate the average tr…

Cited by 24SourcePDFScholar
2021

Equivariant Learning of Stochastic Fields: Gaussian Processes and Steerable Conditional Neural Processes

ICML 2021spotlight

Motivated by objects such as electric fields or fluid streams, we study the problem of learning stochastic fields, i.e. stochastic processes whose samples are fields like those occurring in physics and engineering. Considering general transformations such as rotations and reflections, we show that s…

2021

LieTransformer: Equivariant Self-Attention for Lie Groups

ICML 2021spotlight

Group equivariant neural networks are used as building blocks of group invariant neural networks, which have been shown to improve generalisation performance and data efficiency through principled parameter sharing. Such works have mostly focused on group equivariant convolutions, building on the re…

2021

Neural Ensemble Search for Uncertainty Estimation and Dataset Shift

NeurIPS 2021poster

Ensembles of neural networks achieve superior performance compared to standalone networks in terms of accuracy, uncertainty calibration and robustness to dataset shift. Deep ensembles, a state-of-the-art method for uncertainty estimation, only ensemble random initializations of a fixed architecture.…

2021

Noise Contrastive Meta-Learning for Conditional Density Estimation using Kernel Mean Embeddings

AISTATS 2021poster

Current meta-learning approaches focus on learning functional representations of relationships between variables, \textit{i.e.} estimating conditional expectations in regression. In many applications, however, the conditional distributions cannot be meaningfully summarized solely by expectation (due…

Cited by 14SourcePDFScholar
2021

On Pathologies in KL-Regularized Reinforcement Learning from Expert Demonstrations

NeurIPS 2021poster

KL-regularized reinforcement learning from expert demonstrations has proved successful in improving the sample efficiency of deep reinforcement learning algorithms, allowing them to be applied to challenging physical real-world tasks. However, we show that KL-regularized reinforcement learning with…

2021

Powerpropagation: A sparsity inducing weight reparameterisation

NeurIPS 2021poster

The training of sparse neural networks is becoming an increasingly important tool for reducing the computational footprint of models at training and evaluation, as well enabling the effective scaling up of models. Whereas much work over the years has been dedicated to specialised pruning techniques,…

2021

Vector-valued Gaussian Processes on Riemannian Manifolds via Gauge Independent Projected Kernels

NeurIPS 2021poster

Gaussian processes are machine learning models capable of learning unknown functions in a way that represents uncertainty, thereby facilitating construction of optimal decision-making systems. Motivated by a desire to deploy Gaussian processes in novel areas of science, a rapidly-growing line of res…

Cited by 28SourcePDFScholar
2020

A Unified Stochastic Gradient Approach to Designing Bayesian-Optimal Experiments

AISTATS 2020poster

We introduce a fully stochastic gradient based approach to Bayesian optimal experimental design (BOED). Our approach utilizes variational lower bounds on the expected information gain (EIG) of an experiment that can be simultaneously optimized with respect to both the variational and design paramete…

2020

Bayesian Deep Ensembles via the Neural Tangent Kernel

NeurIPS 2020poster

We explore the link between deep ensembles and Gaussian processes (GPs) through the lens of the Neural Tangent Kernel (NTK): a recent development in understanding the training dynamics of wide neural networks (NNs). Previous work has shown that even in the infinite width limit, when NNs become GPs,…

2020

Bootstrapping neural processes

NeurIPS 2020poster

Unlike in the traditional statistical modeling for which a user typically hand-specify a prior, Neural Processes (NPs) implicitly define a broad class of stochastic processes with neural networks. Given a data stream, NP learns a stochastic process that best describes the data. While this ``data-dri…

2020

Divide, Conquer, and Combine: a New Inference Strategy for Probabilistic Programs with Stochastic Support

ICML 2020poster

Universal probabilistic programming systems (PPSs) provide a powerful framework for specifying rich probabilistic models. They further attempt to automate the process of drawing inferences from these models, but doing this successfully is severely hampered by the wide range of non–standard models th…

Cited by 24SourcePDFScholar
2020

Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise

ICML 2020poster

Stochastic gradient descent with momentum (SGDm) is one of the most popular optimization algorithms in deep learning. While there is a rich theory of SGDm for convex problems, the theory is considerably less developed in the context of deep learning where the problem is non-convex and the gradient n…

2020

Functional Regularisation for Continual Learning with Gaussian Processes

ICLR 2020poster

We introduce a framework for Continual Learning (CL) based on Bayesian inference over the function space rather than the parameters of a deep neural network. This method, referred to as functional regularisation for Continual Learning, avoids forgetting a previous task by constructing and memorising…

Cited by 215SourceScholar
2020

How Robust are the Estimated Effects of Nonpharmaceutical Interventions against COVID-19?

NeurIPS 2020spotlight

To what extent are effectiveness estimates of nonpharmaceutical interventions (NPIs) against COVID-19 influenced by the assumptions our models make? To answer this question, we investigate 2 state-of-the-art NPI effectiveness models and propose 6 variants that make different structural assumptions.…

2020

MetaFun: Meta-Learning with Iterative Functional Updates

ICML 2020poster

We develop a functional encoder-decoder approach to supervised meta-learning, where labeled data is encoded into an infinite-dimensional functional representation rather than a finite-dimensional one. Furthermore, rather than directly producing the representation, we learn a neural update rule resem…

2020

Multiplicative Interactions and Where to Find Them

ICLR 2020poster

We explore the role of multiplicative interaction as a unifying framework to describe a range of classical and modern neural network architectural motifs, such as gating, attention layers, hypernetworks, and dynamic convolutions amongst others. Multiplicative interaction layers as primitive operatio…

Cited by 154SourceScholar
2020

Non-exchangeable feature allocation models with sublinear growth of the feature sizes

AISTATS 2020poster

Feature allocation models are popular models used in different applications such as unsupervised learning or network modeling. In particular, the Indian buffet process is a flexible and simple one-parameter feature allocation model where the number of features grows unboundedly with the number of ob…

Cited by 6SourcePDFScholar
2020

Uncertainty Estimation Using a Single Deep Deterministic Neural Network

ICML 2020poster

We propose a method for training a deterministic deep model that can find and reject out of distribution data points at test time with a single forward pass. Our approach, deterministic uncertainty quantification (DUQ), builds upon ideas of RBF networks. We scale training in these with a novel loss…

2019

A Statistical Approach to Assessing Neural Network Robustness

ICLR 2019poster

We present a new approach to assessing the robustness of neural networks based on estimating the proportion of inputs for which a property is violated. Specifically, we estimate the probability of the event that the property is violated under an input model. Our approach critically varies from the f…

2019

Attentive Neural Processes

ICLR 2019poster

Neural Processes (NPs) (Garnelo et al., 2018) approach regression by learning to map a context set of observed input-output pairs to a distribution over regression functions. Each function models the distribution of the output given an input, conditioned on the context. NPs have the benefit of fitti…

2019

Continual Unsupervised Representation Learning

NeurIPS 2019poster

Continual learning aims to improve the ability of modern learning systems to deal with non-stationary distributions, typically by attempting to learn a series of tasks sequentially. Prior art in the field has largely considered supervised or reinforcement learning tasks, and often assumes full knowl…

2019

Continuous Hierarchical Representations with Poincaré Variational Auto-Encoders

NeurIPS 2019poster

The Variational Auto-Encoder (VAE) is a popular method for learning a generative model and embeddings of the data. Many real datasets are hierarchically structured. However, traditional VAEs map data in a Euclidean latent space which cannot efficiently embed tree-like structures. Hyperbolic spaces…

2019

Disentangling Disentanglement in Variational Autoencoders

ICML 2019oral

We develop a generalisation of disentanglement in variational autoencoders (VAEs)—decomposition of the latent representation—characterising it as the fulfilment of two factors: a) the latent encodings of the data having an appropriate level of overlap, and b) the aggregate encoding of the data confo…

2019

Do Deep Generative Models Know What They Don't Know?

ICLR 2019poster

A neural network deployed in the wild may be asked to make predictions for inputs that were drawn from a different distribution than that of the training data. A plethora of work has demonstrated that it is easy to find or synthesize inputs for which a neural network is highly confident yet wrong.…

Cited by 903SourcePDFScholar
2019

Hybrid Models with Deep and Invertible Features

ICML 2019oral

We propose a neural hybrid model consisting of a linear model defined on a set of features computed by a deep, invertible transformation (i.e. a normalizing flow). An attractive property of our model is that both p(features), the density of the features, and p(targets|features), the predictive distr…

Cited by 110SourcePDFScholar
2019

Information asymmetry in KL-regularized RL

ICLR 2019poster

Many real world tasks exhibit rich structure that is repeated across different parts of the state space or in time. In this work we study the possibility of leveraging such repeated structure to speed up and regularize learning. We start from the KL regularized expected reward objective which introd…

Cited by 109SourcePDFScholar
2019

Neural Probabilistic Motor Primitives for Humanoid Control

ICLR 2019poster

We focus on the problem of learning a single motor module that can flexibly express a range of behaviors for the control of high-dimensional physically simulated humanoids. To do this, we propose a motor architecture that has the general structure of an inverse model with a latent-variable bottlenec…

Cited by 177SourcePDFScholar
2019

Revisiting Reweighted Wake-Sleep for Models with Stochastic Control Flow

UAI 2019poster

Stochastic control-flow models (SCFMs) are a class of generative models that involve branching on choices from discrete random variables. Amortized gradient-based learning of SCFMs is challenging as most approaches targeting discrete variables rely on their continuous relaxations—which can be intrac…

Cited by 54SourcePDFScholar
2019

Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks

ICML 2019oral

Many machine learning tasks such as multiple instance learning, 3D shape recognition, and few-shot image classification are defined on sets of instances. Since solutions to such problems do not depend on the order of elements of the set, models used to address them should be permutation invariant. W…

2019

Variational Bayesian Optimal Experimental Design

NeurIPS 2019spotlight

Bayesian optimal experimental design (BOED) is a principled framework for making efficient use of limited experimental resources. Unfortunately, its applicability is hampered by the difficulty of obtaining accurate estimates of the expected information gain (EIG) of an experiment. To address this, w…

2018

An Analysis of Categorical Distributional Reinforcement Learning

AISTATS 2018poster

Distributional approaches to value-based reinforcement learning model the entire distribution of returns, rather than just their expected values, and have recently been shown to yield state-of-the-art empirical performance. This was demonstrated by the recently proposed C51 algorithm, based on categ…

Cited by 0SourcePDFScholar
2018

Conditional Neural Processes

ICML 2018oral

Deep neural networks excel at function approximation, yet they are typically trained from scratch for each new function. On the other hand, Bayesian methods, such as Gaussian Processes (GPs), exploit prior knowledge to quickly infer the shape of a new function at test time. Yet, GPs are computationa…

Cited by 906SourcePDFScholar
2018

Faithful Inversion of Generative Models for Effective Amortized Inference

NeurIPS 2018poster

Inference amortization methods share information across multiple posterior-inference problems, allowing each to be carried out more efficiently. Generally, they require the inversion of the dependency structure in the generative model, as the modeller must learn a mapping from observations to distri…

Cited by 57SourcePDFScholar
2018

Mix & Match Agent Curricula for Reinforcement Learning

ICML 2018oral

We introduce Mix and match (M&M) – a training framework designed to facilitate rapid and effective learning in RL agents that would be too slow or too challenging to train otherwise.The key innovation is a procedure that allows us to automatically form a curriculum over agents. Through such a curric…

Cited by 96SourcePDFScholar
2018

Modelling sparsity, heterogeneity, reciprocity and community structure in temporal interaction data

NeurIPS 2018poster

We propose a novel class of network models for temporal dyadic interaction data. Our objective is to capture important features often observed in social interactions: sparsity, degree heterogeneity, community structure and reciprocity. We use mutually-exciting Hawkes processes to model the interacti…

2018

Progress & Compress: A scalable framework for continual learning

ICML 2018oral

We introduce a conceptually simple and scalable framework for continual learning domains where tasks are learned sequentially. Our method is constant in the number of parameters and is designed to preserve performance on previously encountered tasks while accelerating learning progress on subsequent…

Cited by 1080SourcePDFScholar
2018

Scaling up the Automatic Statistician: Scalable Structure Discovery using Gaussian Processes

AISTATS 2018poster

Automating statistical modelling is a challenging problem in artificial intelligence. The Automatic Statistician employs a kernel search algorithm using Gaussian Processes (GP) to provide interpretable statistical models for regression problems. However this does not scale due to its O(N^3) running…

Cited by 0SourcePDFScholar
2018

Sequential Attend, Infer, Repeat: Generative Modelling of Moving Objects

NeurIPS 2018spotlight

We present Sequential Attend, Infer, Repeat (SQAIR), an interpretable deep generative model for image sequences. It can reliably discover and track objects through the sequence; it can also conditionally generate future frames, thereby simulating expected motion of objects. This is achieved by expl…

2018

Stochastic Expectation Maximization with Variance Reduction

NeurIPS 2018poster

Expectation-Maximization (EM) is a popular tool for learning latent variable models, but the vanilla batch EM does not scale to large data sets because the whole data set is needed at every E-step. Stochastic Expectation Maximization (sEM) reduces the cost of E-step by stochastic approximation. Howe…

2018

Tighter Variational Bounds are Not Necessarily Better

ICML 2018oral

We provide theoretical and empirical evidence that using tighter evidence lower bounds (ELBOs) can be detrimental to the process of learning an inference network by reducing the signal-to-noise ratio of the gradient estimator. Our results call into question common implicit assumptions that tighter E…

Cited by 246SourcePDFScholar
2017

The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables

ICLR 2017poster

The reparameterization trick enables optimizing large scale stochastic computation graphs via gradient descent. The essence of the trick is to refactor each stochastic node into a differentiable function of its parameters and a random variable with fixed distribution. After refactoring, the gradient…

Cited by 3092SourceScholar
2016

Mondrian Forests for Large-Scale Regression when Uncertainty Matters

AISTATS 2016poster

Many real-world regression problems demand a measure of the uncertainty associated with each prediction. Standard decision forests deliver efficient state-of-the-art predictive performance, but high-quality uncertainty estimates are lacking. Gaussian processes (GPs) deliver uncertainty estimates, b…

Cited by 67SourcePDFScholar