← Search

Sang Bin Moon

1 accepted papers

2024

Optimistic Regret Bounds for Online Learning in Adversarial Markov Decision Processes

UAI 2024poster

The Adversarial Markov Decision Process (AMDP) is a learning framework that deals with unknown and varying tasks in decision-making applications like robotics and recommendation systems. A major limitation of the AMDP formalism, however, is pessimistic regret analysis results in the sense that altho…

Cited by 1SourcePDFScholar