2026
Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates
ICML 2026poster
We study the problem of learning the optimal policy in a discounted, infinite-horizon reinforcement learning (RL) setting in the presence of adversarially corrupted rewards. To address this problem, we develop a novel robust variant of the Q-learning algorithm and analyze it under the challenging as…