← Search

Giorgia Ramponi

20 accepted papers

2026

Fine-tuning Behavioral Cloning Policies with Preference‑Based Reinforcement Learning

ICLR 2026poster

Deploying reinforcement learning (RL) in robotics, industry, and health care is blocked by two obstacles: the difficulty of specifying accurate rewards and the risk of unsafe, data-hungry exploration. We address this by proposing a two-stage framework that first learns a safe initial policy from a r…

Cited by 0SourcecodeScholar
2026

Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference

ICML 2026poster

Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly learn reward functions from heterogeneous feedback types such as demonstrations, comparisons, ratings, rankings, and stops t…

Cited by 0SourceScholar
2026

Multi-agent imitation learning with function approximation: linear Markov games and beyond

ICML 2026poster

In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dynamics and each agent's reward function are linear in some given features. We demonstrate that by leveraging this structure, it is possible to replace t…

Cited by 0SourceScholar
2025

Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning

NeurIPS 2025poster

This paper provides the first expert sample complexity characterization for learning a Nash equilibrium from expert data in Markov Games. We show that a new quantity named the *single policy deviation concentrability coefficient* is unavoidable in the non-interactive imitation learning setting, and…

Cited by 0SourceScholar
2025

Non-rectangular Robust MDPs with Normed Uncertainty Sets

NeurIPS 2025poster

Robust policy evaluation for non-rectangular uncertainty set is generally NP-hard, even in approximation. Consequently, existing approaches suffer from either exponential iteration complexity or significant accuracy gaps. Interestingly, we identify a powerful class of $L_p$-bounded uncertainty sets…

Cited by 0SourceScholar
2025

On the Convergence of Single-Timescale Actor-Critic

NeurIPS 2025poster

We analyze the global convergence of the single-timescale actor-critic (AC) algorithm for the infinite-horizon discounted Markov Decision Processes (MDPs) with finite state spaces. To this end, we introduce an elegant analytical framework for handling complex, coupled recursions inherent in the algo…

Cited by 0SourceScholar
2025

Preference Elicitation for Offline Reinforcement Learning

ICLR 2025poster

Applying reinforcement learning (RL) to real-world problems is often made challenging by the inability to interact with the environment and the difficulty of designing reward functions. Offline RL addresses the first challenge by considering access to an offline dataset of environment interactions l…

Cited by 0SourcePDFScholar
2024

Contextual Bilevel Reinforcement Learning for Incentive Alignment

NeurIPS 2024poster

The optimal policy in various real-world strategic decision-making problems depends both on the environmental configuration and exogenous events. For these settings, we introduce Contextual Bilevel Reinforcement Learning (CB-RL), a stochastic bilevel decision-making model, where the lower level cons…

2024

Exploiting Causal Graph Priors with Posterior Sampling for Reinforcement Learning

ICLR 2024poster

Posterior sampling allows exploitation of prior knowledge on the environment's transition dynamics to improve the sample efficiency of reinforcement learning. The prior is typically specified as a class of parametric distributions, the design of which can be cumbersome in practice, often resulting i…

Cited by 5SourcePDFScholar
2024

Truly No-Regret Learning in Constrained MDPs

ICML 2024spotlight

Constrained Markov decision processes (CMDPs) are a common way to model safety constraints in reinforcement learning. State-of-the-art methods for efficiently solving CMDPs are based on primal-dual algorithms. For these algorithms, all currently known regret bounds allow for *error cancellations* --…

Cited by 12SourcePDFScholar
2023

On Imitation in Mean-field Games

NeurIPS 2023poster

We explore the problem of imitation learning (IL) in the context of mean-field games (MFGs), where the goal is to imitate the behavior of a population of agents following a Nash equilibrium policy according to some unknown payoff function. IL in MFGs presents new challenges compared to single-agent…

Cited by 2SourcePDFScholar
2022

Trust Region Policy Optimization with Optimal Transport Discrepancies: Duality and Algorithm for Continuous Actions

NeurIPS 2022accept

Policy Optimization (PO) algorithms have been proven particularly suited to handle the high-dimensionality of real-world continuous control tasks. In this context, Trust Region Policy Optimization methods represent a popular approach to stabilize the policy updates. These usually rely on the Kullbac…

Cited by 13SourcePDFScholar
2021

Learning in Non-Cooperative Configurable Markov Decision Processes

NeurIPS 2021poster

The Configurable Markov Decision Process framework includes two entities: a Reinforcement Learning agent and a configurator that can modify some environmental parameters to improve the agent's performance. This presupposes that the two actors have the same reward functions. What if the configurator…

Cited by 13SourcePDFScholar
2021

Provably Efficient Learning of Transferable Rewards

ICML 2021spotlight

The reward function is widely accepted as a succinct, robust, and transferable representation of a task. Typical approaches, at the basis of Inverse Reinforcement Learning (IRL), leverage on expert demonstrations to recover a reward function. In this paper, we study the theoretical properties of the…

Cited by 41SourcePDFScholar
2020

Inverse Reinforcement Learning from a Gradient-based Learner

NeurIPS 2020poster

Inverse Reinforcement Learning addresses the problem of inferring an expert's reward function from demonstrations. However, in many applications, we not only have access to the expert's near-optimal behaviour, but we also observe part of her learning process. In this paper, we propose a new algorith…

Cited by 17SourcePDFScholar
2020

Truly Batch Model-Free Inverse Reinforcement Learning about Multiple Intentions

AISTATS 2020poster

We consider Inverse Reinforcement Learning (IRL) about multiple intentions, \ie the problem of estimating the unknown reward functions optimized by a group of experts that demonstrate optimal behaviors. Most of the existing algorithms either require access to a model of the environment or need to re…

Cited by 42SourcePDFScholar