← Search

Alireza Kazemipour

2 accepted papers

2025

Model-Based Exploration in Monitored Markov Decision Processes

ICML 2025poster

A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or malfunctioning, or rewards may be inaccessible during deployment. Monito…

Cited by 1SourcePDFScholar
2024

Beyond Optimism: Exploration With Partially Observable Rewards

NeurIPS 2024poster

Exploration in reinforcement learning (RL) remains an open challenge. RL algorithms rely on observing rewards to train the agent, and if informative rewards are sparse the agent learns slowly or may not learn at all. To improve exploration and reward discovery, popular algorithms rely on optimism.…