← Search

Bruno C. da Silva

5 accepted papers

2026

From Noise to Control: Parameterized Diffusion Policies

ICML 2026poster

We propose Parameterized Diffusion Policy (PDP), a framework that learns a diffusion policy parameterized in a smooth continuous space. By structuring a latent manifold such that distances between latents' values reflect the semantic similarity of physical trajectories, we transform diffusion from a…

Cited by 0SourceScholar
2022

Constrained Offline Policy Optimization

ICML 2022spotlight

In this work we introduce Constrained Offline Policy Optimization (COPO), an offline policy optimization algorithm for learning in MDPs with cost constraints. COPO is built upon a novel offline cost-projection method, which we formally derive and analyze. Our method improves upon the state-of-the-ar…

Cited by 22SourcePDFScholar
2022

Optimistic Linear Support and Successor Features as a Basis for Optimal Policy Transfer

ICML 2022spotlight

In many real-world applications, reinforcement learning (RL) agents might have to solve multiple tasks, each one typically modeled via a reward function. If reward functions are expressed linearly, and the agent has previously learned a set of policies for different tasks, successor features (SFs) c…

2021

Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods

ICML 2021spotlight

Hindsight allows reinforcement learning agents to leverage new observations to make inferences about earlier states and transitions. In this paper, we exploit the idea of hindsight and introduce posterior value functions. Posterior value functions are computed by inferring the posterior distribution…

Cited by 12SourcePDFScholar