← Search

Leif Döring

6 accepted papers

2026

The Role of Target Update Frequencies in Q-Learning

ICML 2026poster

The target network update frequency (TUF) is a central stabilization mechanism in (deep) Q-learning. However, their selection remains poorly understood and is often treated merely as another tunable hyperparameter rather than as a principled design decision. This work provides a theoretical analysis…

Cited by 0SourceScholar
2025

ADDQ: Adaptive distributional double Q-learning

ICML 2025poster

Bias problems in the estimation of Q-values are a well-known obstacle that slows down convergence of Q-learning and actor-critic methods. One of the reasons of the success of modern RL algorithms is partially a direct or indirect overestimation reduction mechanism. We introduce an easy to implement…

2025

Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization

NeurIPS 2025poster

The present article studies the minimization of convex, $L$-smooth functions defined on a separable real Hilbert space. We analyze regularized stochastic gradient descent (reg-SGD), a variant of stochastic gradient descent that uses a Tikhonov regularization with time-dependent, vanishing regulariza…

Cited by 0SourceScholar
2024

Beyond Stationarity: Convergence Analysis of Stochastic Softmax Policy Gradient Methods

ICLR 2024poster

Markov Decision Processes (MDPs) are a formal framework for modeling and solving sequential decision-making problems. In finite time horizons such problems are relevant for instance for optimal stopping or specific supply chain problems, but also in the training of large language models. In contrast…

Cited by 8SourcePDFScholar