← Search

Riccardo Poiani

10 accepted papers

2026

Optimal Rates for Feasible Payoff Set Estimation in Games

ICML 2026spotlight

We study a setting in which two players play a (possibly approximate) Nash equilibrium of a bimatrix game, while a learner observes only their actions and has no knowledge of the equilibrium or the underlying game. A natural question is whether the learner can rationalize the observed behavior by in…

Cited by 0SourceScholar
2024

Optimal Multi-Fidelity Best-Arm Identification

NeurIPS 2024poster

In bandit best-arm identification, an algorithm is tasked with finding the arm with highest mean reward with a specified accuracy as fast as possible. We study multi-fidelity best-arm identification, in which the algorithm can choose to sample an arm at a lower fidelity (less accurate mean estimate)…

Cited by 4SourcePDFScholar
2024

Sub-optimal Experts mitigate Ambiguity in Inverse Reinforcement Learning

NeurIPS 2024poster

Inverse Reinforcement Learning (IRL) deals with the problem of deducing a reward function that explains the behavior of an expert agent who is assumed to act *optimally* in an underlying unknown task. Recent works have studied the IRL problem from the perspective of recovering the *feasible reward s…

Cited by 0SourcePDFScholar
2023

Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive Approach

NeurIPS 2023poster

Policy evaluation via Monte Carlo (MC) simulation is at the core of many MC Reinforcement Learning (RL) algorithms (e.g., policy gradient methods). In this context, the designer of the learning system specifies an interaction budget that the agent usually spends by collecting trajectories of *fixed…

Cited by 1SourcePDFScholar
2023

Truncating Trajectories in Monte Carlo Reinforcement Learning

ICML 2023poster

In Reinforcement Learning (RL), an agent acts in an unknown environment to maximize the expected cumulative discounted sum of an external reward signal, i.e., the expected return. In practice, in many tasks of interest, such as policy optimization, the agent usually spends its interaction budget by…

Cited by 5SourcePDFScholar
2021

Meta-Reinforcement Learning by Tracking Task Non-stationarity

IJCAI 2021poster

Many real-world domains are subject to a structured non-stationarity which affects the agent's goals and the environmental dynamics. Meta-reinforcement learning (RL) has been shown successful for training agents that quickly adapt to related tasks. However, most of the existing meta-RL algorithms fo…

2020

Sequential Transfer in Reinforcement Learning with a Generative Model

ICML 2020poster

We are interested in how to design reinforcement learning agents that provably reduce the sample complexity for learning new tasks by transferring knowledge from previously-solved ones. The availability of solutions to related problems poses a fundamental trade-off: whether to seek policies that are…

Cited by 31SourcePDFScholar