← Search

Marco Pratticò

1 accepted papers

2024

Operator World Models for Reinforcement Learning

NeurIPS 2024poster

Policy Mirror Descent (PMD) is a powerful and theoretically sound methodology for sequential decision-making. However, it is not directly applicable to Reinforcement Learning (RL) due to the inaccessibility of explicit action-value functions. We address this challenge by introducing a novel approach…