← Search

Francesco Emanuele Stradi

11 accepted papers

2025

Data-Dependent Regret Bounds for Constrained MABs

NeurIPS 2025poster

This paper initiates the study of data-dependent regret bounds in constrained MAB settings. These are bounds that depend on the sequence of losses that characterize the problem instance. Thus, in principle they can be much smaller than classical $\widetilde{\mathcal{O}}(\sqrt{T})$ regret bounds, wh…

Cited by 0SourceScholar
2025

Learning Adversarial MDPs with Stochastic Hard Constraints

ICML 2025poster

We study online learning in constrained Markov decision processes (CMDPs) with adversarial losses and stochastic hard constraints, under bandit feedback. We consider three scenarios. In the first one, we address general CMDPs, where we design an algorithm attaining sublinear regret and cumulative po…

Cited by 11SourcePDFScholar
2025

Markov Persuasion Processes: Learning to Persuade From Scratch

NeurIPS 2025poster

In Bayesian persuasion, an informed sender strategically discloses information to a receiver so as to persuade them to undertake desirable actions. Recently, Markov persuasion processes (MPPs) have been introduced to capture sequential scenarios where a sender faces a stream of myopic receivers in a…

Cited by 0SourceScholar
2025

No-Regret Learning Under Adversarial Resource Constraints: A Spending Plan Is All You Need!

NeurIPS 2025poster

We study online decision making problems under resource constraints, where both reward and cost functions are drawn from distributions that may change adversarially over time. We focus on two canonical settings: $(i)$ online resource allocation where rewards and costs are observed before action sele…

Cited by 0SourceScholar
2025

Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization

ICLR 2025poster

We study online learning in constrained MDPs (CMDPs), focusing on the goal of attaining sublinear strong regret and strong cumulative constraint violation. Differently from their standard (weak) counterparts, these metrics do not allow negative terms to compensate positive ones, raising considerable…

Cited by 2SourcePDFScholar
2025

Policy Optimization for CMDPs with Bandit Feedback: Learning Stochastic and Adversarial Constraints

ICML 2025poster

We study online learning in constrained Markov decision processes (CMDPs) in which rewards and constraints may be either stochastic or adversarial. In such settings, stradi et al. (2024) proposed the first best-of-both-worlds algorithm able to seamlessly handle stochastic and adversarial constraints…

Cited by 0SourcePDFScholar
2025

Taming Adversarial Constraints in CMDPs

NeurIPS 2025poster

In constrained MDPs (CMDPs) with adversarial rewards and constraints, a known impossibility result prevents any algorithm from attaining sublinear regret and constraint violation, when competing against a best-in-hindsight policy that satisfies the constraints on average. In this paper, we show how…

Cited by 0SourceScholar
2024

Bandits with Ranking Feedback

NeurIPS 2024poster

In this paper, we introduce a novel variation of multi-armed bandits called bandits with ranking feedback. Unlike traditional bandits, this variation provides feedback to the learner that allows them to rank the arms based on previous pulls, without quantifying numerically the difference in performa…

Cited by 1SourcePDFScholar
2024

Online Learning in CMDPs: Handling Stochastic and Adversarial Constraints

ICML 2024poster

We study online learning in episodic constrained Markov decision processes (CMDPs), where the learner aims at collecting as much reward as possible over the episodes, while satisfying some long-term constraints during the learning process. Rewards and constraints can be selected either stochasticall…

Cited by 6SourcePDFScholar
2024

Online Learning with Off-Policy Feedback in Adversarial MDPs

IJCAI 2024poster

In this paper, we face the challenge of online learning in adversarial Markov decision processes with off-policy feedback. In this setting, the learner chooses a policy, but, differently from the traditional on-policy setting, the environment is explored by means of a different, fixed, and possibly…

Cited by 0SourcePDFScholar
2024

Online Markov Decision Processes Configuration with Continuous Decision Space

AAAI 2024technical

In this paper, we investigate the optimal online configuration of episodic Markov decision processes when the space of the possible configurations is continuous. Specifically, we study the interaction between a learner (referred to as the configurator) and an agent with a fixed, unknown policy, when…

Cited by 10SourcePDFScholar