← Search

Pier Giuseppe Sessa

15 accepted papers

2025

BOND: Aligning LLMs with Best-of-N Distillation

ICLR 2025poster

Reinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models. Yet, a surprisingly simple and strong inference-time strategy is Best-of-N sampling that selects the best generation among N candidates. In this paper, we propose Best-o…

Cited by 26SourcePDFScholar
2025

Optimistic Games for Combinatorial Bayesian Optimization with Application to Protein Design

ICLR 2025poster

Bayesian optimization (BO) is a powerful framework to optimize black-box expensive-to-evaluate functions via sequential interactions. In several important problems (e.g. drug discovery, circuit design, neural architecture search, etc.), though, such functions are defined over large $\textit{combinat…

Cited by 1SourcePDFScholar
2024

Distributionally Robust Model-based Reinforcement Learning with Large State Spaces

AISTATS 2024poster

Three major challenges in reinforcement learning are the complex dynamical systems with large state spaces, the costly data acquisition processes, and the deviation of real-world dynamics from the training environment deployment. To overcome these issues, we study distributionally robust Markov deci…

2024

Group Robust Preference Optimization in Reward-free RLHF

NeurIPS 2024poster

Adapting large language models (LLMs) for specific tasks usually involves fine-tuning through reinforcement learning with human feedback (RLHF) on preference data. While these data often come from diverse labelers' groups (e.g., different demographics, ethnicities, company teams, etc.), traditional…

Cited by 20SourcePDFScholar
2023

How Bad is Selfish Driving? Bounding the Inefficiency of Equilibria in Urban Driving Games

RA-L 2023

We consider the interaction among agents engaging in a driving task and we model it as general-sum game. This class of games exhibits a plurality of different equilibria posing the issue of equilibrium selection. While selecting the most efficient equilibrium (in term of social cost) is often imprac

Cited by 5SourceScholar
2023

Multitask Learning with No Regret: from Improved Confidence Bounds to Active Learning

NeurIPS 2023poster

Multitask learning is a powerful framework that enables one to simultaneously learn multiple related tasks by sharing information between them. Quantifying uncertainty in the estimated tasks is of pivotal importance for many downstream applications, such as online or active learning. In this work, w…

Cited by 4SourcePDFScholar
2022

Efficient Model-based Multi-agent Reinforcement Learning via Optimistic Equilibrium Computation

ICML 2022spotlight

We consider model-based multi-agent reinforcement learning, where the environment transition model is unknown and can only be learned via expensive interactions with the environment. We propose H-MARL (Hallucinated Multi-Agent Reinforcement Learning), a novel sample-efficient algorithm that can effi…

Cited by 19SourcePDFScholar
2022

Movement Penalized Bayesian Optimization with Application to Wind Energy Systems

NeurIPS 2022accept

Contextual Bayesian optimization (CBO) is a powerful framework for sequential decision-making given side information, with important applications, e.g., in wind energy systems. In this setting, the learner receives context (e.g., weather conditions) at each round, and has to choose an action (e.g.,…

Cited by 13SourcePDFScholar
2021

Online Submodular Resource Allocation with Applications to Rebalancing Shared Mobility Systems

ICML 2021spotlight

Motivated by applications in shared mobility, we address the problem of allocating a group of agents to a set of resources to maximize a cumulative welfare objective. We model the welfare obtainable from each resource as a monotone DR-submodular function which is a-priori unknown and can only be lea…

Cited by 3SourcePDFScholar
2020

Contextual Games: Multi-Agent Learning with Side Information

NeurIPS 2020poster

We formulate the novel class of contextual games, a type of repeated games driven by contextual information at each round. By means of kernel-based regularity assumptions, we model the correlation between different contexts and game outcomes and propose a novel online (meta) algorithm that exploits…

Cited by 23SourcePDFScholar
2020

Learning to Play Sequential Games versus Unknown Opponents

NeurIPS 2020poster

We consider a repeated sequential game between a learner, who plays first, and an opponent who responds to the chosen action. We seek to design strategies for the learner to successfully interact with the opponent. While most previous approaches consider known opponent models, we focus on the settin…

Cited by 27SourcePDFScholar
2020

Mixed Strategies for Robust Optimization of Unknown Objectives

AISTATS 2020poster

We consider robust optimization problems, where the goal is to optimize an unknown objective function against the worst-case realization of an uncertain parameter. For this setting, we design a novel sample-efficient algorithm GP-MRO, which sequentially learns about the unknown objective from noisy…

Cited by 18SourcePDFScholar
2019

Bounding Inefficiency of Equilibria in Continuous Actions Games using Submodularity and Curvature

AISTATS 2019poster

Games with continuous strategy sets arise in several machine learning problems (e.g. adversarial learning). For such games, simple no-regret learning algorithms exist in several cases and ensure convergence to coarse correlated equilibria (CCE). The efficiency of such equilibria with respect to a s…

Cited by 14SourcePDFScholar
2019

No-Regret Learning in Unknown Games with Correlated Payoffs

NeurIPS 2019poster

We consider the problem of learning to play a repeated multi-agent game with an unknown reward function. Single player online learning algorithms attain strong regret bounds when provided with full information feedback, which unfortunately is unavailable in many real-world scenarios. Bandit feedback…

Cited by 48SourcePDFScholar