2023
Belief Projection-Based Reinforcement Learning for Environments with Delayed Feedback
NeurIPS 2023poster
We present a novel actor-critic algorithm for an environment with delayed feedback, which addresses the state-space explosion problem of conventional approaches. Conventional approaches use an augmented state constructed from the last observed state and actions executed since visiting the last obser…