← Search

Bianca Marin Moreno

2 accepted papers

2025

Online Episodic Convex Reinforcement Learning

ICML 2025poster

We study online learning in episodic finite-horizon Markov decision processes (MDPs) with convex objective functions, known as the concave utility reinforcement learning (CURL) problem. This setting generalizes RL from linear to convex losses on the state-action distribution induced by the agent’s p…

Cited by 0SourcePDFScholar
2024

MetaCURL: Non-stationary Concave Utility Reinforcement Learning

NeurIPS 2024poster

We explore online learning in episodic loop-free Markov decision processes on non-stationary environments (changing losses and probability transitions). Our focus is on the Concave Utility Reinforcement Learning problem (CURL), an extension of classical RL for handling convex performance criteria in…

Cited by 0SourcePDFScholar