← Search

Orestis Papadigenopoulos

7 accepted papers

2024

Contextual Pandora’s Box

AAAI 2024technical

Pandora’s Box is a fundamental stochastic optimization problem, where the decision-maker must find a good alternative, while minimizing the search cost of exploring the value of each alternative. In the original formulation, it is assumed that accurate distributions are given for the values of all t…

Cited by 6SourcePDFScholar
2023

Last Switch Dependent Bandits with Monotone Payoff Functions

ICML 2023poster

In a recent work, Laforgue et al. introduce the model of last switch dependent (LSD) bandits, in an attempt to capture nonstationary phenomena induced by the interaction between the player and the environment. Examples include satiation, where consecutive plays of the same action lead to decreased p…

Cited by 4SourcePDFScholar
2022

Non-Stationary Bandits under Recharging Payoffs: Improved Planning with Sublinear Regret

NeurIPS 2022accept

The stochastic multi-armed bandit setting has been recently studied in the non-stationary regime, where the mean payoff of each action is a non-decreasing function of the number of rounds passed since it was last played. This model captures natural behavioral aspects of the users which crucially det…

Cited by 4SourcePDFScholar
2021

Combinatorial Blocking Bandits with Stochastic Delays

ICML 2021spotlight

Recent work has considered natural variations of the {\em multi-armed bandit} problem, where the reward distribution of each arm is a special function of the time passed since its last pulling. In this direction, a simple (yet widely applicable) model is that of {\em blocking bandits}, where an arm…

Cited by 16SourcePDFScholar
2021

Contextual Blocking Bandits

AISTATS 2021poster

We study a novel variant of the multi-armed bandit problem, where at each time step, the player observes an independently sampled context that determines the arms’ mean rewards. However, playing an arm blocks it (across all contexts) for a fixed number of future time steps. The above contextual sett…

Cited by 28SourcePDFScholar
2021

Recurrent Submodular Welfare and Matroid Blocking Semi-Bandits

NeurIPS 2021poster

A recent line of research focuses on the study of stochastic multi-armed bandits (MAB), in the case where temporal correlations of specific structure are imposed between the player's actions and the reward distributions of the arms. These correlations lead to (sub-)optimal solutions that exhibit int…

Cited by 11SourcePDFScholar