2024
Mitigating Partial Observability in Sequential Decision Processes via the Lambda Discrepancy
NeurIPS 2024poster
Reinforcement learning algorithms typically rely on the assumption that the environment dynamics and value function can be expressed in terms of a Markovian state representation. However, when state information is only partially observable, how can an agent learn such a state representation, and how…