2026
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
ICML 2026poster
Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift (thus slow recovery), and leaving unanswered the principled question of how …