2026
Adaptive Reinforcement Learning for Unobservable Random Delays
ICML 2026poster
In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov decision process (MDP), which assumes that the agent observes the system state instantaneously, selects an action without delay, and executes it immediately. In re…