← Search

Alberto Sardinha

6 accepted papers

2026

Centralized Training with Hybrid Execution in Multi-Agent Reinforcement Learning via Predictive Observation Imputation (Abstract Reprint)

AAAI 2026technical

We study hybrid execution in multi-agent reinforcement learning (MARL), a paradigm where agents aim to complete cooperative tasks with arbitrary communication levels at execution time by taking advantage of information-sharing among the agents. Under hybrid execution, the communication level can ran

Cited by 0SourcePDFScholar
2026

Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning

ICLR 2026poster

In this work, we contribute the first approach to solve infinite-horizon discounted general-utility Markov decision processes (GUMDPs) in the single-trial regime, i.e., when the agent's performance is evaluated based on a single trajectory. First, we provide some fundamental results regarding policy…

Cited by 0SourceScholar
2025

The Number of Trials Matters in Infinite-Horizon General-Utility Markov Decision Processes

ICML 2025spotlight

The general-utility Markov decision processes (GUMDPs) framework generalizes the MDPs framework by considering objective functions that depend on the frequency of visitation of state-action pairs induced by a given policy. In this work, we contribute with the first analysis on the impact of the numb…

Cited by 0SourcePDFScholar
2024

TEAMSTER: Model-Based Reinforcement Learning for Ad Hoc Teamwork (Abstract Reprint)

AAAI 2024technical

This paper investigates the use of model-based reinforcement learning in the context of ad hoc teamwork. We introduce a novel approach, named TEAMSTER, where we propose learning both the environment's model and the model of the teammates' behavior separately. Compared to the state-of-the-art PLASTIC…

Cited by 0SourcePDFScholar
2022

Perceive, Represent, Generate: Translating Multimodal Information to Robotic Motion Trajectories

IROS 2022poster

We present Perceive-Represent-Generate (PRG), a novel three-stage framework that maps perceptual information of different modalities (e.g., visual or sound), corresponding to a series of instructions, to a sequence of movements to be executed by a robot. In the first stage, we perceive and preproces…

Cited by 1SourceScholar
2021

Online Learning Meets Machine Translation Evaluation: Finding the Best Systems with the Least Human Effort

ACL 2021long

In Machine Translation, assessing the quality of a large amount of automatic translations can be challenging. Automatic metrics are not reliable when it comes to high performing systems. In addition, resorting to human evaluators can be expensive, especially when evaluating multiple systems. To over…