← Search

Francesco Belardinelli

12 accepted papers

2026

Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning

AAAI 2026technical

Many reinforcement learning algorithms, particularly those that rely on return estimates for policy improvement, can suffer from poor sample efficiency and training instability due to high-variance return estimates. In this paper we leverage new results from off-policy evaluation; it has recently be

Cited by 0SourcePDFScholar
2025

Probabilistic Shielding for Safe Reinforcement Learning

AAAI 2025technical

In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximize their reward, must often also behave in a safe manner, including at training time. Thus, much attention in recent years has been given to Safe RL, where an agent aims to learn an optimal policy among all policies that sat…

Cited by 0SourcePDFScholar
2024

Stability of Multi-Agent Learning in Competitive Networks: Delaying the Onset of Chaos

AAAI 2024technical

The behaviour of multi agent learning in competitive network games is often studied within the context of zero sum games, in which convergence guarantees may be obtained. However, outside of this class the behaviour of learning is known to display complex behaviours and convergence cannot be always…

Cited by 2SourcePDFScholar
2023

Automatically Verifying Expressive Epistemic Properties of Programs

AAAI 2023technical

We propose a new approach to the verification of epistemic properties of programmes. First, we introduce the new ``program-epistemic'' logic L_PK, which is strictly richer and more general than similar formalisms appearing in the literature. To solve the verification problem in an efficient way, we…

2023

Beyond Strict Competition: Approximate Convergence of Multi-agent Q-Learning Dynamics

IJCAI 2023poster

The behaviour of multi-agent learning in competitive settings is often considered under the restrictive assumption of a zero-sum game. Only under this strict requirement is the behaviour of learning well understood; beyond this, learning dynamics can often display non-convergent behaviours which pre…

Cited by 2SourcePDFScholar
2023

Honesty Is the Best Policy: Defining and Mitigating AI Deception

NeurIPS 2023spotlight

Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with language models, the goal of being evaluated as truthful). There are a number of e…

Cited by 35SourcePDFScholar
2023

Scalable Verification of Strategy Logic through Three-Valued Abstraction

IJCAI 2023poster

The model checking problem for multi-agent systems against Strategy Logic specifications is known to be non-elementary. On this logic several fragments have been defined to tackle this issue but at the expense of expressiveness. In this paper, we propose a three-valued semantics for Strategy Logic u…

2023

The Impact of Exploration on Convergence and Performance of Multi-Agent Q-Learning Dynamics

ICML 2023poster

Understanding the impact of exploration on the behaviour of multi-agent learning has, so far, benefited from the restriction to potential, or network zero-sum games in which convergence to an equilibrium can be shown. Outside of these classes, learning dynamics rarely converge and little is known ab…

Cited by 2SourcePDFScholar
2022

In a Nutshell, the Human Asked for This: Latent Goals for Following Temporal Specifications

ICLR 2022poster

We address the problem of building agents whose goal is to learn to execute out-of distribution (OOD) multi-task instructions expressed in temporal logic (TL) by using deep reinforcement learning (DRL). Recent works provided evidence that the agent's neural architecture is a key feature when DRL age…

2021

Reasoning About Agents That May Know Other Agents’ Strategies

IJCAI 2021poster

We study the semantics of knowledge in strategic reasoning. Most existing works either implicitly assume that agents do not know one another’s strategies, or that all strategies are known to all; and some works present inconsistent mixes of both features. We put forward a novel semantics for Strateg…

Cited by 10SourcePDFScholar