← Search

Moritz Akiya Zanger

5 accepted papers

2026

Contextual Similarity Distillation: Ensemble Uncertainties with a Single Model

ICLR 2026poster

Uncertainty quantification is a critical aspect of reinforcement learning and deep learning, with numerous applications ranging from efficient exploration and stable offline reinforcement learning to outlier detection in medical diagnostics. The scale of modern neural networks, however, complicates…

Cited by 0SourceScholar
2026

Universal Value-Function Uncertainties

ICLR 2026poster

Estimating epistemic uncertainty in value functions is a crucial challenge for many aspects of reinforcement learning (RL), including efficient exploration, safe decision-making, and offline RL. While deep ensembles provide a robust method for quantifying value uncertainty, they come with significan…

Cited by 0SourceScholar
2025

How Ensembles of Distilled Policies Improve Generalisation in Reinforcement Learning

NeurIPS 2025poster

In the zero-shot policy transfer setting in reinforcement learning, the goal is to train an agent on a fixed set of training environments so that it can generalise to similar, but unseen, testing environments. Previous work has shown that policy distillation after training can sometimes produce a po…

Cited by 0SourceScholar
2025

Value Improved Actor Critic Algorithms

NeurIPS 2025poster

To learn approximately optimal acting policies for decision problems, modern Actor Critic algorithms rely on deep Neural Networks (DNNs) to parameterize the acting policy and greedification operators to iteratively improve it. The reliance on DNNs suggests an improvement that is gradient based, whi…

Cited by 0SourceScholar
2024

Diverse Projection Ensembles for Distributional Reinforcement Learning

ICLR 2024poster

In contrast to classical reinforcement learning, distributional RL algorithms aim to learn the distribution of returns rather than their expected value. Since the nature of the return distribution is generally unknown a priori or arbitrarily complex, a common approach finds approximations within a s…

Cited by 3SourcePDFScholar