← Search

Dirk Van Der Hoeven

12 accepted papers

2025

When Lower-Order Terms Dominate: Adaptive Expert Algorithms for Heavy-Tailed Losses

NeurIPS 2025poster

We consider the problem setting of prediction with expert advice with possibly heavy-tailed losses, i.e.\ the only assumption on the losses is an upper bound on their second moments, denoted by $\theta$. We develop adaptive algorithms that do not require any prior knowledge about the range or the se…

Cited by 0SourceScholar
2023

Delayed Bandits: When Do Intermediate Observations Help?

ICML 2023poster

We study a $K$-armed bandit with delayed feedback and intermediate observations. We consider a model, where intermediate observations have a form of a finite state, which is observed immediately after taking an action, whereas the loss is observed after an adversarially chosen delay. We show that th…

Cited by 2SourcePDFScholar
2023

Nonstochastic Contextual Combinatorial Bandits

AISTATS 2023poster

We study a contextual version of online combinatorial optimisation with full and semi-bandit feedback. In this sequential decision-making problem, an online learner has to select an action from a combinatorial decision space after seeing a vector-valued context in each round. As a result of its acti…

Cited by 6SourcePDFScholar
2023

Trading-Off Payments and Accuracy in Online Classification with Paid Stochastic Experts

ICML 2023poster

We investigate online classification with paid stochastic experts. Here, before making their prediction, each expert must be paid. The amount that we pay each expert directly influences the accuracy of their prediction through some unknown Lipschitz ``productivity'' function. In each round, the lear…

Cited by 0SourcePDFScholar
2022

A Near-Optimal Best-of-Both-Worlds Algorithm for Online Learning with Feedback Graphs

NeurIPS 2022accept

We consider online learning with feedback graphs, a sequential decision-making framework where the learner's feedback is determined by a directed graph over the action set. We present a computationally-efficient algorithm for learning in this framework that simultaneously achieves near-optimal regre…

Cited by 22SourcePDFScholar
2022

Learning on the Edge: Online Learning with Stochastic Feedback Graphs

NeurIPS 2022accept

The framework of feedback graphs is a generalization of sequential decision-making with bandit or full information feedback. In this work, we study an extension where the directed feedback graph is stochastic, following a distribution similar to the classical Erdős-Rényi model. Specifically, in each…

Cited by 16SourcePDFScholar
2021

Beyond Bandit Feedback in Online Multiclass Classification

NeurIPS 2021poster

We study the problem of online multiclass classification in a setting where the learner's feedback is determined by an arbitrary directed graph. While including bandit feedback as a special case, feedback graphs allow a much richer set of applications, including filtering and label efficient classif…

Cited by 14SourcePDFScholar