← Search

Dabeen Lee

8 accepted papers

2026

Neural Logistic Bandits

ICML 2026poster

We study the problem of neural logistic bandits, where the main task is to learn an unknown reward function within a logistic link function using a neural network. Existing approaches either exhibit unfavorable dependencies on $\kappa$, where $1/\kappa$ represents the minimum variance of reward dist…

Cited by 0SourceScholar
2025

An Optimistic Algorithm for online CMDPS with Anytime Adversarial Constraints

ICML 2025poster

Online safe reinforcement learning (RL) plays a key role in dynamic environments, with applications in autonomous driving, robotics, and cybersecurity. The objective is to learn optimal policies that maximize rewards while satisfying safety constraints modeled by constrained Markov decision processe…

Cited by 0SourcePDFScholar
2025

Infinite-Horizon Reinforcement Learning with Multinomial Logit Function Approximation

AISTATS 2025poster

We study model-based reinforcement learning with non-linear function approximation where the transition function of the underlying Markov decision process (MDP) is given by a multinomial logit (MNL) model. We develop a provably efficient discounted value iteration-based algorithm that works for both…

Cited by 0SourceScholar
2025

Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span

AISTATS 2025poster

This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear mixture Markov decision processes (MDPs) under the Bellman optimality condition. Our algorithm for linear mixture MDPs achieves a nearly minimax optimal regret upper bound of $\widetilde{\ma…

Cited by 0SourceScholar
2025

Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs

AISTATS 2025poster

We study the problem of infinite-horizon average-reward reinforcement learning with linear Markov decision processes (MDPs). The associated Bellman operator of the problem not being a contraction makes the algorithm design challenging. Previous approaches either suffer from computational inefficienc…

Cited by 0SourceScholar