← Search

Marco Mussi

17 accepted papers

2026

Reusing Trajectories in Policy Gradients Enables Fast Convergence

ICML 2026poster

*Policy gradient* (PG) methods are a class of effective *reinforcement learning* algorithms, particularly when dealing with continuous control problems. They rely on fresh *on-policy* data, making them sample-inefficient and requiring $\mathcal{O}(\epsilon^{-2})$ trajectories to reach an $\epsilon$-…

Cited by 0SourceScholar
2025

Convergence Analysis of Policy Gradient Methods with Dynamic Stochasticity

ICML 2025poster

*Policy gradient* (PG) methods are effective *reinforcement learning* (RL) approaches, particularly for continuous problems. While they optimize stochastic (hyper)policies via action- or parameter-space exploration, real-world applications often require deterministic policies. Existing PG convergenc…

Cited by 0SourcePDFScholar
2025

Position: Constants are Critical in Regret Bounds for Reinforcement Learning

ICML 2025poster

Mainstream research in theoretical RL is currently focused on designing online learning algorithms with regret bounds that match the corresponding regret lower bound up to multiplicative constants (and, sometimes, logarithmic terms). In this position paper, we constructively question this trend, arg…

Cited by 0SourcePDFScholar
2025

Tightening Regret Lower and Upper Bounds in Restless Rising Bandits

NeurIPS 2025poster

*Restless* Multi-Armed Bandits (MABs) are a general framework designed to handle real-world decision-making problems where the expected rewards evolve over time, such as in recommender systems and dynamic pricing. In this work, we investigate from a theoretical standpoint two well-known structured s…

Cited by 0SourceScholar
2025

Towards Theoretical Understanding of Sequential Decision Making with Preference Feedback

ICML 2025poster

The success of sequential decision-making approaches, such as *reinforcement learning* (RL), is closely tied to the availability of a reward feedback. However, designing a reward function that encodes the desired objective is a challenging task. In this work, we address a more realistic scenario: se…

Cited by 0SourcePDFScholar
2024

Autoregressive Bandits

AISTATS 2024poster

Autoregressive processes naturally arise in a large variety of real-world scenarios, including stock markets, sales forecasting, weather prediction, advertising, and pricing. When facing a sequential decision-making problem in such a context, the temporal dependence between consecutive observations…

2024

Best Arm Identification for Stochastic Rising Bandits

ICML 2024spotlight

Stochastic Rising Bandits (SRBs) model sequential decision-making problems in which the expected reward of the available options increases every time they are selected. This setting captures a wide range of scenarios in which the available options are learning entities whose performance improves (in…

2024

Factored-Reward Bandits with Intermediate Observations

ICML 2024poster

In several real-world sequential decision problems, at every step, the learner is required to select different actions. Every action affects a specific part of the system and generates an observable intermediate effect. In this paper, we introduce the Factored-Reward Bandits (FRBs), a novel setting…

Cited by 1SourcePDFScholar
2024

Graph-Triggered Rising Bandits

ICML 2024poster

In this paper, we propose a novel generalization of rested and restless bandits where the evolution of the arms' expected rewards is governed by a graph defined over the arms. An edge connecting a pair of arms $(i,j)$ represents the fact that a pull of arm $i$ *triggers* the evolution of arm $j$, an…

Cited by 4SourcePDFScholar
2024

Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning

NeurIPS 2024poster

*Constrained Reinforcement Learning* (CRL) tackles sequential decision-making problems where agents are required to achieve goals by maximizing the expected return while meeting domain-specific constraints, which are often formulated on expected costs. In this setting, *policy-based* methods are wid…

Cited by 2SourcePDFScholar
2024

Learning Optimal Deterministic Policies with Stochastic Policy Gradients

ICML 2024spotlight

Policy gradient (PG) methods are successful approaches to deal with continuous reinforcement learning (RL) problems. They learn stochastic parametric (hyper)policies by either exploring in the space of actions or in the space of parameters. Stochastic controllers, however, are often undesirable from…

Cited by 2SourcePDFScholar
2023

Dynamic Pricing with Volume Discounts in Online Settings

AAAI 2023technical

According to the main international reports, more pervasive industrial and business-process automation, thanks to machine learning and advanced analytic tools, will unlock more than 14 trillion USD worldwide annually by 2030. In the specific case of pricing problems, which constitute the class of pr…

Cited by 7SourcePDFScholar