Enhancing Value Function Estimation through First-Order State-Action Dynamics in Offline Reinforcement Learning
In offline reinforcement learning (RL), updating the value function with the discrete-time Bellman Equation often encounters challenges due to the limited scope of available data. This limitation stems from the Bellman Equation, which cannot accurately predict the value of unvisited states. To addre…