2024
RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning
ICML 2024poster
In this paper, we propose an off-policy deep reinforcement learning (DRL) method utilizing the average reward criterion. While most existing DRL methods employ the discounted reward criterion, this can potentially lead to a discrepancy between the training objective and performance metrics in contin…