← Search

Lorenzo Mancini

2 accepted papers

2025

Federated UCBVI: Communication-Efficient Federated Regret Minimization with Heterogeneous Agents

AISTATS 2025poster

In this paper, we present the Federated Upper Confidence Bound Value Iteration algorithm ($\texttt{Fed-UCBVI}$), a novel extension of the $\texttt{UCBVI}$ algorithm (Azar et al., 2017) tailored for the federated learning framework. We prove that the regret of $\texttt{Fed-UCBVI}$ scales as $\tilde O…

Cited by 0SourceScholar
2025

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean Field Games

ICML 2025poster

We introduce Mean Field Trust Region Policy Optimization (MF-TRPO), a novel algorithm designed to compute approximate Nash equilibria for ergodic Mean Field Games (MFGs) in finite state-action spaces. Building on the well-established performance of TRPO in the reinforcement learning (RL) setting, we…

Cited by 0SourcePDFScholar